跳到正文
原文
arXiv · Human-Computer Interaction· Jinyi Ye, Scott Counts, Gaurav Verma, Kate Lytvynets, Weiwei Yang·· 3 小时前AI 评分50

从图像到任务:真实场景中多模态 LLM 交互的特征分析

From Images to Tasks: Characterizing Multimodal LLM Interactions in the Wild

AI 导读

研究分析 Microsoft Copilot 中 40,000+ 张图像上传对话,提出涵盖感知、认知与生成十种能力的分层框架。多数图像上传任务涉及多能力组合,多模态使用覆盖比纯文本更广、更依赖跨模态对齐。253 个现有基准对感知与推理覆盖集中,文本、代码与数据生成等常见工作流测试不足,该结论在独立 ChatGPT 数据集上得到验证。

来源:arXiv · Human-Computer Interaction · arxiv.org