How rankings work · methodology 1.4.1
Overall rank is pure capability — the Power Index. CursorBench is weighted heavily among a small core of hard benchmarks. Missing cores are imputed from priors so sparse cherry-pick scores cannot leapfrog dense frontier models. Speed and price never enter the rank.
Every core family always contributes: direct evidence when present, otherwise a discounted prior impute (not zero, not pure drop-and-renormalize). Reliability rises only with real coverage. A small finishing bonus applies only for real CursorBench rows—not substitutes—so models missing agent suites cannot invent coding-agent strength.
Everything else in the benchmark directory remains available for inspection and per-benchmark leaderboards. Only this suite feeds overall Power.
| Family | Weight | Role |
|---|---|---|
| CursorBench | 20% | Real coding-agent work. If missing, we substitute a discounted blend of SWE / Terminal / live coding. |
| SWE-Bench Pro | 14% | Repository engineering (falls back to Verified / Rebench). |
| Terminal-Bench | 10% | Agentic terminal tasks (falls back to Hard / long-horizon). |
| Humanity's Last Exam | 10% | Hard general academic intelligence. |
| GPQA Diamond | 8% | Expert science reasoning. |
| GDPval-AA | 8% | Real economic deliverable tasks. |
| LiveCodeBench / SciCode | 7% | Fresh competitive and scientific code. |
| MCP Atlas / τ-bench | 7% | Tool and MCP agent use. |
| ARC-AGI-2 | 6% | Novel abstract reasoning. |
| Hard math (AIME / FrontierMath) | 5% | Competition math ceiling. |
| MMMU-Pro | 5% | Multimodal (small so vision specialists cannot dominate power). |
Output speed, time-to-first-token, list price, and task cost never enter the overall rank. They remain filterable columns so you can still buy for budget or latency. For pure power, sort by rank.
Commercial relationships, advertising, and sponsorship never change weights, verification, or ranks. Changelog records methodology bumps.