Coding
Weighted toward coding capability and tool use, for developer-focused use cases.
| Rank | Model | Score | Confidence | Evidence | Coding | Tool Use | Reasoning |
|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet · Anthropic | 87.4 | 84% | 2,251 | 86 | 100 | 100 |
| 2 | GPT Frontier · OpenAI | 86.1 | 84% | 2,251 | 100 | 88 | 87 |
| 3 | Gemini Pro · Google DeepMind | 72.8 | 84% | 2,251 | 64 | 80 | 73 |
| 4 | Llama Open · Meta AI | 8.0 | 72% | 446 | 0 | 0 | 0 |
Scores are based on sample data for demonstration purposes.
Formula version 1.0.0 · weights v1.0.0 · calculated 2026-09-21. See the methodology page for how scores, confidence and evidence are combined.