Skip to content
Hot eventLive

研究:LLM骨干模型演进对相关性评估的影响

1 reports1 sources4 hr ago updated

Get the story

AI overview

2026年10月8日,arXiv(Information Retrieval)发布一手研究,考察在固定提示词下,使用UMBRELA与EXAM两类提示对Gemini、GPT、Qwen、Llama连续版本的相关性判断表现。研究未发现新版本必然更贴近人类判断的一致证据;聚合分数相近或提升也不代表判断稳定,早期版本的正确判断可能在后续版本丢失。研究提醒,为某一骨干版本设计验证的提示词,不能假定在同族新模型上等效或更优。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 8, 2026
  1. arXiv · Information Retrieval
    骨干模型演进对 LLM 相关性评估的影响

    研究在固定提示词下,用 UMBRELA 与 EXAM 两类提示评估 Gemini、GPT、Qwen、Llama 连续版本的相关性判断表现。未发现新版本必然更贴近人类判断的一致证据,聚合分数相近或提升也不代表判断稳定,早期版本的正确判断可能在后续版本丢失。研究提醒,为某一骨干版本设计验证的提示词,不能假定在同族新模型上等效或更优。

Heat trend

Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –

02.557.510Oct8Oct8Oct8Oct8

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.