VisAudit:多模态智能体可视化诊断与修复基准
Get the story
2026年10月5日,arXiv Machine Learning Theory 发布一手报道,介绍 VisAudit 基准,用于评测多模态智能体的可视化诊断、修复与验证能力。该基准覆盖 21 种图表类型、10 类缺陷,包含 1,900 个有缺陷实例和 300 个初始正确图表,并设置诊断修复、自主修复与开放世界验证三条赛道。基准通过受控扰动构建,注入的缺陷可定义且可恢复。实验显示,最强评测模型在自主修复设定下仅完全恢复 47.4% 的有缺陷图表。目前未见更早或相互矛盾的报道。
Generated from reports · updated 15 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Machine Learning TheoryVisAudit:评测多模态智能体的视觉诊断与修复能力
VisAudit 是一个评测可视化诊断、修复与验证的基准,覆盖 21 种图表类型、10 类缺陷,含 1,900 个有缺陷实例和 300 个初始正确图表,并设诊断修复、自主修复与开放世界验证三条赛道。基准通过受控扰动构建,注入缺陷可定义且可恢复。实验显示,最强评测模型在自主修复设定下仅完全恢复 47.4% 的有缺陷图表。
Heat trend
Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.