Benchmarks

Each benchmark below documents its own methodology, dataset and limitations — scores are only comparable within the same benchmark.

  • Code Generation Eval

    Sample data

    Code generation

    Metric: pass@1 (%) · Version 1.2 · Updated 2026-08-01

  • Cost & Latency Index

    Sample data

    Operational efficiency

    Metric: composite index (lower is better) · Version 1.0 · Updated 2026-09-01

  • General Reasoning Eval

    Sample data

    Multi-step reasoning

    Metric: accuracy (%) · Version 1.0 · Updated 2026-08-15