Skip to content
Back to leaderboard

How ranking works

Understand the evidence and method behind a ranking built from public evaluations.

0—100

Shared evidence, a clear order.

Combine public evaluations while minimizing conflicts with known scores. The overall and four category boards show up to 30 models.

The consensus index converts evidence differences supporting the original ranking to 0-100 for comparison. It is not accuracy or a percentage of capability difference. Close support can tie; the complete evidence still determines rank.

  1. 01

    First confirm it is the same model

    Normalize names and versions across boards. Choose one representative configuration for each model with a fixed rule rather than the highest score. Exclude anonymous test codes and scores produced with mixed models; officially public Preview versions are eligible.

  2. 02

    Compare models that were actually tested

    Compare real scores from the same evaluation only. When both sides publish uncertainty, small differences within the error range are treated closer to a tie. No measurement means no comparison.

  3. 03

    Set weights before combining evidence

    Each evaluation uses a predetermined budget. The overall board requires at least three organizations, three evidence families and three tracks; repeated captures or display slices from one source do not add voting weight.

  4. 04

    Find the complete ranking with the fewest conflicts

    Evaluations can disagree. We choose the order that violates the least net voting weight from shared evidence, then calculate the display index separately. Missing scores are not filled with zero or guessed.

A BALANCED VIEW

General tests, human preference and specialist evaluations together.

General evaluations account for 30%, human preference for 10%, and specialist evaluations for 60%. New evaluations are checked for overlap before receiving a share; missing shares are not transferred to other evaluations.

  • 综合评测30%AA Index
  • 真人盲选10%Arena Text · Arena 创作盲选
  • 编程与设计12%Arena WebDev · DeepSWE v1.1 · TapTap Maker · LiveBench · 编程综合
  • 写作与表达9%LiveBench · 语言与指令 · Creative Writing v3 · Longform Writing
  • 数学与推理12%LiveBench · 推理与数学 · FrontierMath v2 · Tiers 1–3 · FrontierMath v2 · Tier 4 · Chess Puzzles · Mystery Game Puzzles
  • 知识与事实6%SimpleQA Verified · GPQA Diamond
  • 视觉理解6%Arena Vision
  • 工具与办公6%APEX-Agents 1.1 · τ³-Banking
  • 中文与多语言6%AA 中文/多语言
  • 行业专业任务3%Vals Finance Agent

Coding, reasoning, knowledge and professional work use separate real evaluations and are calculated independently. Vision and multilingual evidence remain in the overall board; renaming categories does not add voting weight. Overall score is not an arithmetic average of category scores. A category usually needs at least two valid evaluations and five comparable models to appear. Knowledge currently uses two Epoch evaluations with the same-organization limitation stated. Creative preference and web development evidence remain in the overall board; the first version does not create separate aesthetics or writing boards.

You may also want to know

Why can a model with less evidence still appear on the board?
The amount of evaluation is different from capability. The overall board requires three organizations and cross-capability evidence. Most categories require two organizations; knowledge uses two evaluations from one organization with the limitation stated. Missing scores are not counted as zero, though they can still introduce bias.
What does evidence-sensitive mean?
A ranking is marked evidence-sensitive when removing an evaluation or organization, changing a weight by 20%, or changing error handling moves the rank by three or more places, removes eligibility in some scenarios, or leaves a comparison incomplete. The range is not a 95% confidence interval and does not include unknown scores.
Does a higher-ranked model always win pairwise comparisons?
Not necessarily. A can beat B, B beat C, and C beat A. The complete board balances these conflicts, so non-adjacent ranks can differ from pairwise comparisons. The mathematical optimum minimizes total conflict under the current rules; it does not prove the real-world order of capability.
Why might the ranking differ from my experience?
The board aggregates public evaluations and reflects the capabilities supported by that evidence. Your experience also depends on product version, reasoning mode, tools and long-task stability. We use new evaluations to test old ranks; close scores should not be read as a clear strength difference.
When will a new model appear?
HOTAI checks upstream results four times a day. A new model enters the ranking only after an evaluator publishes a score; a page refresh, price update or our capture time is not a new evaluation.
What happens when a board is temporarily unavailable?
Valid source snapshots remain usable. Within the same evaluation version, temporarily missing rows can use records verified within the last seven days. Valid scores still present on a public board are not removed just because values stay unchanged. Official withdrawals, corrections or version changes invalidate or replace them; a failed run keeps the previous valid board and timestamp.
Does the consensus index double-count specialist evaluations?
We check overlap using public datasets and methodology notes, and cap the total share of related sources; unknown correlations remain a limitation. The AA index currently has 30%; Arena overall text and the creative track share the original 10% human-preference budget at 5% each. Their votes overlap, so they remain one evidence family rather than two independent evaluations.
Do price or speed affect the ranking?
No. Price is shown only to explain API cost, consistently per million tokens with the vendor source. It does not represent subscription fees.
View calculation details and the current version

Method 2026.09-public-consensus-v15. Weighted incomplete Kemeny ranking is used. Net support M is aggregated for each jointly evaluated pair; the objective minimizes reversed net support. Integer optimization returns the optimum and bounds, and only fully validated results are published.

When both sides publish standard errors, net support is 2Φ(score difference / combined standard error) − 1 with zero covariance by default; other comparisons use only the raw lead. Unknown error is not zero error, and ordinal treatment of small differences remains a limitation. Voting weight applies to potential model pairs, so sources covering more models use more comparison positions; nominal budgets are not exact contributions to final rank.

The display index preserves the source order: it reverses adjacent models, allows other models to reorder, and computes the least added reverse net support. These non-negative supports are accumulated and mapped with a sigmoid against fixed anchors to 0-100. Equally optimal alternatives retain zero spacing; one decimal is shown without an artificial minimum gap. The index does not drive ranking; the board set, anchors and evidence still affect it, and a score gap is not a capability distance. Ties use fixed model IDs; integer optimization uses HiGHS 1.15.3, and normal CDF conversion matches SciPy norm.cdf. Sources, protocols, eligibility and each run's inputs are versioned. No cross-component order is published without a shared evidence network.

Fixed anchor models: claude-fable-5、gpt-5-6-sol、kimi-k-3、qwen-3-8-max、gpt-5-4、claude-opus-4-8、claude-sonnet-5、grok-4-5、gemini-3-5-flash、glm-5-2、claude-sonnet-4-6、qwen-3-7-max、kimi-k-2-6、deepseek-v-4-pro、qwen-3-6-plus、minimax-m-3、gpt-5-4-mini、grok-4-3. Anchors define the index scale and do not prescribe vendor order; indices from different categories are not directly comparable.

View every evaluation →