Back to overall leaderboardCAPABILITY PROFILE BEHIND THE SCORE
GPT-5 2025-08-07
OpenAI · 2025-08-07 released · 10/01 14:05 updated
Overall consensus index—Not ranked overall
Category scores1/ 4 categories
Recorded scores0evaluations
Context window—Token
API input / output · per million tokens¥8.38 / ¥67.05Cached ¥0.84List $1.25 / $10Official vendor price
See where each model excels.
Each capability is scored separately; insufficient evidence is left blank.
编程More evidence needed—Not enough comparable scores推理More evidence needed—Not enough comparable scores知识#2256.32 evaluations专业办公More evidence needed—Not enough comparable scores
Scores reflect ranking support within each reference group and cannot be added or compared as absolute capability levels.
Every score has a source.
Publicly reported scores for this model. Expand a row to see configuration and usage.
Scored evaluations without a result
These evaluations have not published a score for this model; missing scores are not treated as zero.
Arena TextGPQA DiamondChess PuzzlesCreative Writing v3Longform WritingArena VisionArena WebDevDeepSWE v1.1TapTap MakerMystery Game PuzzlesSimpleQA VerifiedLiveBench · 编程综合LiveBench · 语言与指令FrontierMath v2 · Tiers 1–3Vals Finance AgentLiveBench · 推理与数学Arena 创作盲选FrontierMath v2 · Tier 4APEX-Agents 1.1AA Index
Some things remain unknown
Missing evaluations are not counted as zero. Ranks can change with new evidence; close scores should not be overinterpreted.
See the methodology →