By Google DeepMind
Multimodal model emphasizing long-context and tool use.
Official rankings are evidence-weighted and profile-specific. Community opinion is shown separately below.
See where the model lands under each ranking profile.
Sample profile used to demonstrate multi-provider ranking comparisons.
Results below come from different benchmarks with different tasks and metrics — they are not directly comparable to each other, only across models within the same benchmark.
| Benchmark | Score | Sample size | Confidence | Evidence | Evaluated |
|---|---|---|---|---|---|
| General Reasoning Eval | 82 % | 500 | 85% | L2Benchmark Result | 2026-08-15 |
| Code Generation Eval | 75 % | 300 | 80% | L2Benchmark Result | 2026-08-01 |
No reviews match this filter yet.