Methodology
How we score and compare models.
We never self-report quality. Every score comes from an independent, public source, and every price is our live price list next to OpenRouter's public list price.
Human preference — LMArena
LMArena (formerly Chatbot Arena) shows people two anonymous answers to the same prompt and asks which is better. Millions of these blind votes are turned into an Elo-style rating. We show:
- Arena Elo — the overall text leaderboard
- Arena Coding, Math and Hard Prompts — the same leaderboard filtered to those prompt categories
- WebDev Arena — people vote on web apps that each model builds
Data comes from the public lmarena-ai/leaderboard-dataset. A gap of about 20 Elo points or less is usually within the confidence interval, so treat close scores as ties.
Capability evals — Epoch AI
The Epoch AI Benchmarking Hub runs standard evaluations itself and collects results from other leaderboards. We use:
- GPQA Diamond — PhD-level science questions
- SWE-bench Verified — real GitHub issues fixed end to end
- Terminal-Bench 2 — agent tasks in a real terminal (best published agent harness)
- AIME (OTIS mock 2024–2025) and FrontierMath tiers 1–3 — competition and research-level maths
- ARC-AGI-2 — abstract reasoning puzzles (ARC Prize)
- Humanity's Last Exam and SimpleQA Verified — expert knowledge and factual accuracy
- Epoch Capabilities Index (ECI) — Epoch's composite across many benchmarks
Epoch AI data is licensed under CC BY 4.0. Source: Epoch AI, ‘Capabilities & benchmarking’, epoch.ai/benchmarks.
Contamination-resistant — LiveBench
LiveBench releases new questions regularly and grades them automatically, without an LLM judge. We average task scores within each category and then average the categories to get an overall score from 0 to 100. We keep every past release for trend charts.
Matching model names
Each source names models differently (claude-opus-5-5_max, claude-opus-5.5-high). We normalise names by lowercasing, treating dots and dashes in version numbers the same, and removing vendor prefixes, release dates and reasoning-effort suffixes. A hand-checked alias list covers the rest. Ambiguous names are left unmatched rather than guessed.
When a lab publishes several reasoning-effort variants of one model (low, high, max…), we show the best-scoring variant for each benchmark. The compare view shows which variant a score came from when you hover it.
Price — Qubax vs OpenRouter
The Qubax price is what you pay per 1M input and output tokens, taken from our live price list (the same prices on model pages and in billing). The OpenRouter price is OpenRouter's public list price for the same model. For sorting we use a blended price with a 3:1 input-to-output ratio:
blended = (input × 3 + output) / 4
Latency and success rate
Time to first token (TTFT) is measured from sending the request to receiving the first streamed token. It combines real traffic through Qubax with a small synthetic probe (“Reply with exactly: pong”) sent every 15 minutes, so models with little traffic still have fresh numbers. We show p50 and p95 over the last 7 days. Probes are never billed.
Success rate is the share of requests that returned a usable answer. Errors caused by the caller — closed connections, insufficient credits, rate limits and invalid parameters — are excluded.
How often data updates
| Source | We check | Source publishes |
|---|---|---|
| LMArena | Every 6 hours | Weekly or more often |
| Epoch AI | Every 6 hours | As new runs finish |
| LiveBench | Every 6 hours | Periodic releases |
| Latency probes | Every 15 minutes | — |
| Prices | Live | — |
The page itself is cached for up to 10 minutes.
What we don't show
- Our costs or margins
- Any individual user's requests — reliability numbers are aggregates
- Vendor-reported benchmark numbers — only independent, public sources