arXiv · Computers and Society· Jason Miklian, Kristian Hoelscher, John E. Katsos·· 2 小时前精选AI 评分62
随机鹦鹉还是和谐共唱?测试五种主流大语言模型复制人类调查的能力
Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data
AI 导读
论文比较了420名硅谷程序员真实调查回答与ChatGPT Thinking 5 Pro、Claude Sonnet 4.5 Pro plus Claude CoWork 1.123、Gemini Advanced 2.5 Pro、Incredible 1.0和DeepSeek 3.2生成的合成数据。
推荐理由
研究对比五种主流大语言模型生成合成调查数据与真实人类回答的差异,揭示当前模型只能复现常规共识而无法产出反直觉发现。
来源:arXiv · Computers and Society · arxiv.org