General Reasoning Eval
Sample dataMulti-step reasoning
Definition & methodology
- Methodology
- Sample methodology description: multiple-choice and free-response reasoning tasks scored for correctness.
- Dataset
- Sample dataset of reasoning problems spanning math, logic and common sense.
- Metric
- accuracy (%) (higher is better)
- Version
- 1.0
- Last updated
- 2026-08-15
- Limitations
- Sample limitation note: dataset size is small and may not generalize across all reasoning task types.
Model results
| Model | Score | Sample size | Evidence | Evaluated |
|---|---|---|---|---|
| Claude Sonnet | 86 % | 500 | L2Benchmark Result | 2026-08-15 |
| GPT Frontier | 84 % | 500 | L2Benchmark Result | 2026-08-15 |
| Gemini Pro | 82 % | 500 | L2Benchmark Result | 2026-08-15 |
| Llama Open | 71 % | 40 | L2Benchmark Result | 2026-07-01 |