By Meta AI
Open-weight model included to demonstrate coverage of open models.
Official rankings are evidence-weighted and profile-specific. Community opinion is shown separately below.
See where the model lands under each ranking profile.
Sample profile with fewer benchmark entries to demonstrate low-evidence handling.
Results below come from different benchmarks with different tasks and metrics — they are not directly comparable to each other, only across models within the same benchmark.
| Benchmark | Score | Sample size | Confidence | Evidence | Evaluated |
|---|---|---|---|---|---|
| General Reasoning Eval | 71 % | 40 | 35% | L2Benchmark Result | 2026-07-01 |
Sample review: not as strong on complex reasoning tasks but the cost makes it viable for high-volume, lower-stakes use cases.