Reasoning
Weighted toward multi-step reasoning and knowledge depth.
| Rank | Model | Score | Confidence | Evidence | Reasoning | Knowledge | Instruction Following |
|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet · Anthropic | 93.7 | 100% | 1,230 | 100 | 76 | 100 |
| 2 | GPT Frontier · OpenAI | 85.5 | 100% | 1,230 | 87 | 88 | 81 |
| 3 | Gemini Pro · Google DeepMind | 78.2 | 100% | 1,230 | 73 | 100 | 69 |
| 4 | Llama Open · Meta AI | 7.0 | 85% | 220 | 0 | 0 | 0 |
Scores are based on sample data for demonstration purposes.
Formula version 1.0.0 · weights v1.0.0 · calculated 2026-09-21. See the methodology page for how scores, confidence and evidence are combined.