Hot eventLive
研究:LLM骨干模型演进对相关性评估的影响
1 reports1 sources4 hr ago updated
Get the story
AI overview
2026年10月8日,arXiv(Information Retrieval)发布一手研究,考察在固定提示词下,使用UMBRELA与EXAM两类提示对Gemini、GPT、Qwen、Llama连续版本的相关性判断表现。研究未发现新版本必然更贴近人类判断的一致证据;聚合分数相近或提升也不代表判断稳定,早期版本的正确判断可能在后续版本丢失。研究提醒,为某一骨干版本设计验证的提示词,不能假定在同族新模型上等效或更优。
Generated from reports · updated 3 hr ago
LatestOct 8
arXiv研究称未发现新版本必然更贴近人类判断,早期正确判断可能在后续版本丢失。Timeline
Follow the coverage from different angles.
Oct 8, 2026
- arXiv · Information Retrieval骨干模型演进对 LLM 相关性评估的影响
研究在固定提示词下,用 UMBRELA 与 EXAM 两类提示评估 Gemini、GPT、Qwen、Llama 连续版本的相关性判断表现。未发现新版本必然更贴近人类判断的一致证据,聚合分数相近或提升也不代表判断稳定,早期版本的正确判断可能在后续版本丢失。研究提醒,为某一骨干版本设计验证的提示词,不能假定在同族新模型上等效或更优。
Heat trend
Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.