By OpenAI
General-purpose frontier model with strong multimodal capabilities.
Official rankings are evidence-weighted and profile-specific. Community opinion is shown separately below.
See where the model lands under each ranking profile.
Sample profile used to demonstrate comparison and ranking flows across providers.
Results below come from different benchmarks with different tasks and metrics — they are not directly comparable to each other, only across models within the same benchmark.
| Benchmark | Score | Sample size | Confidence | Evidence | Evaluated |
|---|---|---|---|---|---|
| General Reasoning Eval | 84 % | 500 | 85% | L2Benchmark Result | 2026-08-15 |
| Code Generation Eval | 80 % | 300 | 80% | L2Benchmark Result | 2026-08-01 |
| Cost & Latency Index | 48 index | — | 60% | L1Source | 2026-09-01 |
Independent qualitative assessments, distinct from automated benchmark results.
Independent evaluator (sample)
Sample expert note: excellent multimodal handling, occasional verbosity in long responses.
Score: 85/100
No reviews match this filter yet.