How ranking works
Understand the evidence and method behind a ranking built from public evaluations.
0—100
Shared evidence, a clear order.
Combine public evaluations while minimizing conflicts with known scores. The overall and four category boards show up to 30 models.
The consensus index converts evidence differences supporting the original ranking to 0-100 for comparison. It is not accuracy or a percentage of capability difference. Close support can tie; the complete evidence still determines rank.
- 01
First confirm it is the same model
Normalize names and versions across boards. Choose one representative configuration for each model with a fixed rule rather than the highest score. Exclude anonymous test codes and scores produced with mixed models; officially public Preview versions are eligible.
- 02
Compare models that were actually tested
Compare real scores from the same evaluation only. When both sides publish uncertainty, small differences within the error range are treated closer to a tie. No measurement means no comparison.
- 03
Set weights before combining evidence
Each evaluation uses a predetermined budget. The overall board requires at least three organizations, three evidence families and three tracks; repeated captures or display slices from one source do not add voting weight.
- 04
Find the complete ranking with the fewest conflicts
Evaluations can disagree. We choose the order that violates the least net voting weight from shared evidence, then calculate the display index separately. Missing scores are not filled with zero or guessed.
General tests, human preference and specialist evaluations together.
General evaluations account for 30%, human preference for 10%, and specialist evaluations for 60%. New evaluations are checked for overlap before receiving a share; missing shares are not transferred to other evaluations.
- 综合评测30%AA Index
- 真人盲选10%Arena Text · Arena 创作盲选
- 编程与设计12%Arena WebDev · DeepSWE v1.1 · TapTap Maker · LiveBench · 编程综合
- 写作与表达9%LiveBench · 语言与指令 · Creative Writing v3 · Longform Writing
- 数学与推理12%LiveBench · 推理与数学 · FrontierMath v2 · Tiers 1–3 · FrontierMath v2 · Tier 4 · Chess Puzzles · Mystery Game Puzzles
- 知识与事实6%SimpleQA Verified · GPQA Diamond
- 视觉理解6%Arena Vision
- 工具与办公6%APEX-Agents 1.1 · τ³-Banking
- 中文与多语言6%AA 中文/多语言
- 行业专业任务3%Vals Finance Agent
Coding, reasoning, knowledge and professional work use separate real evaluations and are calculated independently. Vision and multilingual evidence remain in the overall board; renaming categories does not add voting weight. Overall score is not an arithmetic average of category scores. A category usually needs at least two valid evaluations and five comparable models to appear. Knowledge currently uses two Epoch evaluations with the same-organization limitation stated. Creative preference and web development evidence remain in the overall board; the first version does not create separate aesthetics or writing boards.
You may also want to know
Why can a model with less evidence still appear on the board?
What does evidence-sensitive mean?
Does a higher-ranked model always win pairwise comparisons?
Why might the ranking differ from my experience?
When will a new model appear?
What happens when a board is temporarily unavailable?
Does the consensus index double-count specialist evaluations?
Do price or speed affect the ranking?
View calculation details and the current version
Method 2026.09-public-consensus-v15. Weighted incomplete Kemeny ranking is used. Net support M is aggregated for each jointly evaluated pair; the objective minimizes reversed net support. Integer optimization returns the optimum and bounds, and only fully validated results are published.
When both sides publish standard errors, net support is 2Φ(score difference / combined standard error) − 1 with zero covariance by default; other comparisons use only the raw lead. Unknown error is not zero error, and ordinal treatment of small differences remains a limitation. Voting weight applies to potential model pairs, so sources covering more models use more comparison positions; nominal budgets are not exact contributions to final rank.
The display index preserves the source order: it reverses adjacent models, allows other models to reorder, and computes the least added reverse net support. These non-negative supports are accumulated and mapped with a sigmoid against fixed anchors to 0-100. Equally optimal alternatives retain zero spacing; one decimal is shown without an artificial minimum gap. The index does not drive ranking; the board set, anchors and evidence still affect it, and a score gap is not a capability distance. Ties use fixed model IDs; integer optimization uses HiGHS 1.15.3, and normal CDF conversion matches SciPy norm.cdf. Sources, protocols, eligibility and each run's inputs are versioned. No cross-component order is published without a shared evidence network.
Fixed anchor models: claude-fable-5、gpt-5-6-sol、kimi-k-3、qwen-3-8-max、gpt-5-4、claude-opus-4-8、claude-sonnet-5、grok-4-5、gemini-3-5-flash、glm-5-2、claude-sonnet-4-6、qwen-3-7-max、kimi-k-2-6、deepseek-v-4-pro、qwen-3-6-plus、minimax-m-3、gpt-5-4-mini、grok-4-3. Anchors define the index scale and do not prescribe vendor order; indices from different categories are not directly comparable.