用合作原则评估视觉语言模型的VQA表现
Get the story
2026年10月5日,arXiv(Computation and Language,一手)发表研究《基于合作原则评估视觉语言模型的 VQA 性能》。研究让 VLM 生成违反 Grice 准则的问题修饰语(包括非必要、歧义或虚假信息),并据此评测 ChatGPT、Claude、Gemini 和 Llava 的 VQA 表现,发现违规存在时这些模型表现下降。实验进一步对比人类与 VLM 的语用推理差异:人类处理 VLM 引发的违规时认知负荷更低,但 VLM 在此类情况下准确率反而更差。
Generated from reports · updated 16 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Computation and Language基于合作原则评估视觉语言模型的 VQA 性能
研究用 VLM 生成违反 Grice 准则的问题修饰语(非必要、歧义或虚假信息),评测 ChatGPT、Claude、Gemini 和 Llava 的 VQA 性能,发现违规存在时这些模型表现下降。实验进一步对比人类与 VLM 的语用推理差异:人类处理 VLM 引发的违规认知负荷更低,但 VLM 在此类情况下准确率反而更差。
Heat trend
Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.