Skip to content
Back to overall leaderboard

o3 2025-04-16

OpenAI · 2025-04-16 released · 10/01 14:05 updated

Overall consensus index0.8Outside the overall top 30
Category scores1/ 4 categories
Recorded scores11evaluations
Context window—Token
API input / output · per million tokens¥13.41 / ¥53.64Cached ¥3.35List $2 / $8Official vendor price
CAPABILITY PROFILE

See where each model excels.

Each capability is scored separately; insufficient evidence is left blank.

Scores reflect ranking support within each reference group and cannot be added or compared as absolute capability levels.

UNDERSTANDING THE RANK

How stable is the overall rank?

Remove an evaluation or organization, adjust weights and error handling, then observe the rank.

Rank after eligibility checks64-77 Low confidence

1 scenarios lack enough evidence. With the original candidate set it is #75-78. This is not a confidence interval and excludes unpublished scores.

BEHIND THE SCORE

Every score has a source.

Publicly reported scores for this model. Expand a row to see configuration and usage.

综合评测与体验5 items

推理3 items

知识2 items

视觉理解1 items

Scored evaluations without a result

These evaluations have not published a score for this model; missing scores are not treated as zero.

Some things remain unknown

Missing evaluations are not counted as zero. Ranks can change with new evidence; close scores should not be overinterpreted.

See the methodology →