Hot eventLive
五款LLM合成数据复现人类调查对比研究
1 reports1 sources3 hr ago updated
Get the story
AI overview
2026-10-02 arXiv 发表 Computers and Society 论文,比较五种主流大语言模型复现人类调查的能力。研究以420名硅谷程序员的真实调查回答为基准,测试 ChatGPT Thinking 5 Pro、Claude Sonnet 4.5 Pro plus Claude CoWork 1.123、Gemini Advanced 2.5 Pro、Incredible 1.0 和 DeepSeek 3.2 生成的合成数据与真实回答的匹配程度。论文标题为“随机鹦鹉还是和谐共唱?”,旨在评估 LLM 在模拟人类调查响应方面的表现。
Generated from reports · updated 2 hr ago
Timeline
Follow the coverage from different angles.
Oct 2, 2026
- arXiv · Computers and Society精选随机鹦鹉还是和谐共唱?测试五种主流大语言模型复制人类调查的能力
论文比较了420名硅谷程序员真实调查回答与ChatGPT Thinking 5 Pro、Claude Sonnet 4.5 Pro plus Claude CoWork 1.123、Gemini Advanced 2.5 Pro、Incredible 1.0和DeepSeek 3.2生成的合成数据。
Heat trend
There is not enough continuous observation data to draw a trend yet.