Back to Benchmarks

Methodology

How we score and compare models.

We never self-report quality. Every score comes from an independent, public source, and every price is our live price list next to OpenRouter's public list price.

Human preference — LMArena

LMArena (formerly Chatbot Arena) shows people two anonymous answers to the same prompt and asks which is better. Millions of these blind votes are turned into an Elo-style rating. We show:

  • Arena Elo — the overall text leaderboard
  • Arena Coding, Math and Hard Prompts — the same leaderboard filtered to those prompt categories
  • WebDev Arena — people vote on web apps that each model builds

Data comes from the public lmarena-ai/leaderboard-dataset. A gap of about 20 Elo points or less is usually within the confidence interval, so treat close scores as ties.

Capability evals — Epoch AI

The Epoch AI Benchmarking Hub runs standard evaluations itself and collects results from other leaderboards. We use:

  • GPQA Diamond — PhD-level science questions
  • SWE-bench Verified — real GitHub issues fixed end to end
  • Terminal-Bench 2 — agent tasks in a real terminal (best published agent harness)
  • AIME (OTIS mock 2024–2025) and FrontierMath tiers 1–3 — competition and research-level maths
  • ARC-AGI-2 — abstract reasoning puzzles (ARC Prize)
  • Humanity's Last Exam and SimpleQA Verified — expert knowledge and factual accuracy
  • Epoch Capabilities Index (ECI) — Epoch's composite across many benchmarks

Epoch AI data is licensed under CC BY 4.0. Source: Epoch AI, ‘Capabilities & benchmarking’, epoch.ai/benchmarks.

Contamination-resistant — LiveBench

LiveBench releases new questions regularly and grades them automatically, without an LLM judge. We average task scores within each category and then average the categories to get an overall score from 0 to 100. We keep every past release for trend charts.

Matching model names

Each source names models differently (claude-opus-5-5_max, claude-opus-5.5-high). We normalise names by lowercasing, treating dots and dashes in version numbers the same, and removing vendor prefixes, release dates and reasoning-effort suffixes. A hand-checked alias list covers the rest. Ambiguous names are left unmatched rather than guessed.

When a lab publishes several reasoning-effort variants of one model (low, high, max…), we show the best-scoring variant for each benchmark. The compare view shows which variant a score came from when you hover it.

Price — Qubax vs OpenRouter

The Qubax price is what you pay per 1M input and output tokens, taken from our live price list (the same prices on model pages and in billing). The OpenRouter price is OpenRouter's public list price for the same model. For sorting we use a blended price with a 3:1 input-to-output ratio:

blended = (input × 3 + output) / 4

Latency and success rate

Time to first token (TTFT) is measured from sending the request to receiving the first streamed token. It combines real traffic through Qubax with a small synthetic probe (“Reply with exactly: pong”) sent every 15 minutes, so models with little traffic still have fresh numbers. We show p50 and p95 over the last 7 days. Probes are never billed.

Success rate is the share of requests that returned a usable answer. Errors caused by the caller — closed connections, insufficient credits, rate limits and invalid parameters — are excluded.

How often data updates

SourceWe checkSource publishes
LMArenaEvery 6 hoursWeekly or more often
Epoch AIEvery 6 hoursAs new runs finish
LiveBenchEvery 6 hoursPeriodic releases
Latency probesEvery 15 minutes—
PricesLive—

The page itself is cached for up to 10 minutes.

What we don't show

  • Our costs or margins
  • Any individual user's requests — reliability numbers are aggregates
  • Vendor-reported benchmark numbers — only independent, public sources