General Reasoning Eval

Sample data

Multi-step reasoning

Definition & methodology

Methodology
Sample methodology description: multiple-choice and free-response reasoning tasks scored for correctness.
Dataset
Sample dataset of reasoning problems spanning math, logic and common sense.
Metric
accuracy (%) (higher is better)
Version
1.0
Last updated
2026-08-15
Limitations
Sample limitation note: dataset size is small and may not generalize across all reasoning task types.

Model results

Results by model for General Reasoning Eval
ModelScoreSample sizeEvidenceEvaluated
Claude Sonnet86 %500L2Benchmark Result2026-08-15
GPT Frontier84 %500L2Benchmark Result2026-08-15
Gemini Pro82 %500L2Benchmark Result2026-08-15
Llama Open71 %40L2Benchmark Result2026-07-01