Skip to content
ModelRankAI
  • Models
  • Rankings
  • Benchmarks
  • Compare
  • Methodology
  • Reviews
  • About

ModelRankAI

Evidence-driven AI model evaluation.

Models

  • Models
  • Compare

Rankings

  • Rankings
  • Benchmarks

About

  • Methodology
  • About
  • Reviews

© 2026 Artificial Intelligence DataBase. All rights reserved.

Admin
Models / Claude Sonnet

Claude Sonnet

Sample data

By Anthropic

A balanced general-purpose model tuned for reasoning and coding tasks.

  • Status: active
  • Context: 200,000 tokens
  • Modalities: text, code, image
  • Released: 2026-02-01
Add to comparisonBrowse rankings
Performance snapshot

How this model ranks

Official rankings are evidence-weighted and profile-specific. Community opinion is shown separately below.

82.6top available profile

Ranking positions

See where the model lands under each ranking profile.

General AI Reasoning#1
82.6

80% confidence · 2,882 evidence

Coding#1
87.4

84% confidence · 2,251 evidence

Reasoning#1
93.7

100% confidence · 1,230 evidence

Value / Quality#2
60.2

80% confidence · 2,051 evidence

Multi-dimensional performance

Reasoning100
100% evidence confidence500 evidence
Instruction Following100
100% evidence confidence250 evidence
Writing100
100% evidence confidence80 evidence
Tool Use100
100% evidence confidence200 evidence
Coding86
100% evidence confidence300 evidence
Knowledge76
100% evidence confidence400 evidence
Speed73
100% evidence confidence1,000 evidence
Context6
71% evidence confidence1 evidence
Multimodal0
100% evidence confidence150 evidence
Cost Efficiency0
71% evidence confidence1 evidenceStale

About

Sample profile used to demonstrate model pages, dimension scoring and benchmark evidence display.

Evidence

Benchmark results

Results below come from different benchmarks with different tasks and metrics — they are not directly comparable to each other, only across models within the same benchmark.

Benchmark results for this model
BenchmarkScoreSample sizeConfidenceEvidenceEvaluated
General Reasoning Eval86 %50085%L2Benchmark Result2026-08-15
Code Generation Eval78 %30080%L2Benchmark Result2026-08-01
Cost & Latency Index42 index—60%L1Source2026-09-01

Expert evaluations

Independent qualitative assessments, distinct from automated benchmark results.

  • Sample Reviewer AL3Expert Review

    Independent evaluator (sample)

    Sample expert note: strong structured reasoning and reliable instruction-following in extended testing.

    Score: 88/100

Community reviews

★★★★★3.6(2 reviews)

Based on a small number of reviews — shown as a conservative estimate rather than the raw average (4.5) until more reviews come in.

  • 5 star1
  • 4 star1
  • 3 star0
  • 2 star0
  • 1 star0

Community ratings are a separate signal from expert evaluations and benchmark evidence — see the model’s Evidence section above for those.

All usesCodingWritingResearchCustomer supportImage generationGeneral useOther

No reviews match this filter yet.

Write a review

Rating