Hot eventLive
OMIT基准:测量哲学分歧下的框架不变遗漏偏差
1 reports1 sources3 hr ago updated
Get the story
AI overview
研究者推出 OMIT 评测基准,用于测量哲学分歧下的框架不变遗漏偏差。该基准包含 218 组配对框架场景、覆盖 10 类冲突,由 LLM 驱动的五视角哲学人格面板构建,五视角分别为功利主义、义务论、美德伦理、关怀伦理与契约论。对 8 个 LLM 的评估显示,遗漏偏差普遍存在,但在同一模型族内与模型规模呈负相关;在 4 种推理时干预中,先让模型考虑道德原则再作答,可降低遗漏偏差并提升框架一致性。
Generated from reports · updated 3 hr ago
LatestOct 7
OMIT基准发布:218组配对场景、10类冲突,测得遗漏偏差普遍存在且与模型规模负相关。Timeline
Follow the coverage from different angles.
Oct 7, 2026
- arXiv · Computation and LanguageOMIT:测量哲学分歧下框架不变的遗漏偏差
研究者推出 OMIT 评测基准,包含 218 组配对框架场景、覆盖 10 类冲突,由 LLM 驱动的五视角哲学人格面板(功利主义、义务论、美德伦理、关怀伦理、契约论)构建。对 8 个 LLM 的评估显示遗漏偏差普遍存在,但在同一模型族内与模型规模呈负相关;4 种推理时干预中,先让模型考虑道德原则再作答可降低遗漏偏差并提升框架一致性。
Heat trend
Current heat 9·Comparable peak 10(Oct 7)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.