Multi-dimensional
Reasoning, coding, knowledge, speed, cost, context and more are scored independently.
Compare models across reasoning, coding, cost and reliability — backed by transparent, versioned evidence.
Sample data — for demonstration purposes only.
Separate evidence, benchmark performance and real-world community opinion instead of hiding everything behind one unexplained score.
Reasoning, coding, knowledge, speed, cost, context and more are scored independently.
Every official signal can be traced to a benchmark, expert review or internal evaluation.
Real user ratings are useful context, but remain visibly separate from official rankings.
Anthropic
A balanced general-purpose model tuned for reasoning and coding tasks.
OpenAI
General-purpose frontier model with strong multimodal capabilities.
Google DeepMind
Multimodal model emphasizing long-context and tool use.
Meta AI
Open-weight model included to demonstrate coverage of open models.
Sample review: handled a multi-file refactor with clear explanations at each step. Minor slowdowns on very long contexts.
coding · 2026-08-20Sample review: image understanding was noticeably better than the previous version we used for support tickets with screenshots.
support · 2026-08-12Sample review: good tone control for long-form writing, occasionally over-explains simple points.
writing · 2026-08-05Sample review: dropped an entire research paper corpus in and got coherent synthesis across the whole set.
research · 2026-07-28Use a focused ranking profile or build your own side-by-side comparison.