arXiv · Computation and Language· Yihuai Hong, Shauli Ravfogel, Chen Zhao, Eunsol Choi·· 7 hr agoSelectedAI score60
Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment
Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment
AI brief
论文提出 CoT-Interpretability Alignment(CIA)指标,衡量大语言模型链式推理轨迹与内部推理策略的一致性。在两跳问答、提示干预和整数乘法三个任务上评估三个模型,对齐度仅 44.8-75.9%。通过后训练同时优化任务准确率和参数忠实度信号,可在保持或提升准确率的同时显著改善 CoT 参数忠实度,并提供泛化模式分析。代码与数据已公开。
Why it matters
提出了衡量链式推理与内部计算一致性的 CIA 指标,为审计大模型推理可信性提供了可复现框架。
Source: arXiv · Computation and Language · arxiv.org