Skip to content
Hot eventLive

论文提出干预迁移法评估LLM评分准则

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026年10月9日,arXiv机器学习理论栏目发布论文(一手),提出干预迁移(IT)方法,用于评估大语言模型生成的评分准则:若两条准则在响应被扰动为通过或失败其中一条时同步变化,则判定二者相似。论文以HealthBench为案例,用Qwen3.8-27B、Deepseek-V4-Flash与Opus-5生成的准则评估GPT-5.6-Terra的响应。结果显示,降低响应质量的扰动可迁移到另一准则的更低分,但提升扰动不能可靠迁移到更高分,呈现不对称性。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 9, 2026
  1. arXiv · Machine Learning Theory
    用干预迁移评估评分准则生成

    论文提出干预迁移(IT)方法评估 LLM 生成的评分准则:若两条准则在响应被扰动为通过/失败其中一条时同步变化,则二者相似。在 HealthBench 案例中,Qwen3.8-27B、Deepseek-V4-Flash 与 Opus-5 生成的准则用于评估 GPT-5.6-Terra 的响应;降低响应的扰动可迁移到另一准则的更低分,但提升扰动不能可靠迁移到更高分,呈现不对称性。

Heat trend

Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –

02.557.510Oct9Oct9Oct9Oct9

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.