By Anthropic
A balanced general-purpose model tuned for reasoning and coding tasks.
Official rankings are evidence-weighted and profile-specific. Community opinion is shown separately below.
See where the model lands under each ranking profile.
Sample profile used to demonstrate model pages, dimension scoring and benchmark evidence display.
Results below come from different benchmarks with different tasks and metrics — they are not directly comparable to each other, only across models within the same benchmark.
| Benchmark | Score | Sample size | Confidence | Evidence | Evaluated |
|---|---|---|---|---|---|
| General Reasoning Eval | 86 % | 500 | 85% | L2Benchmark Result | 2026-08-15 |
| Code Generation Eval | 78 % | 300 | 80% | L2Benchmark Result | 2026-08-01 |
| Cost & Latency Index | 42 index | — | 60% | L1Source | 2026-09-01 |
Independent qualitative assessments, distinct from automated benchmark results.
Independent evaluator (sample)
Sample expert note: strong structured reasoning and reliable instruction-following in extended testing.
Score: 88/100
Sample review: good tone control for long-form writing, occasionally over-explains simple points.