Grok 4.6
xAI · 2026-08-12 released · 10/01 14:05 updated
See where each model excels.
Each capability is scored separately; insufficient evidence is left blank.
Scores reflect ranking support within each reference group and cannot be added or compared as absolute capability levels.
How stable is the overall rank?
Remove an evaluation or organization, adjust weights and error handling, then observe the rank.
All completed comparisons retained eligibility. With the original candidate set it is #16-29. This is not a confidence interval and excludes unpublished scores.
Every score has a source.
Publicly reported scores for this model. Expand a row to see configuration and usage.
综合评测与体验5 items
编程3 items
推理5 items
知识2 items
专业办公2 items
视觉理解1 items
Scored evaluations without a result
These evaluations have not published a score for this model; missing scores are not treated as zero.
Some things remain unknown
Missing evaluations are not counted as zero. Ranks can change with new evidence; close scores should not be overinterpreted.
See the methodology →