arXiv · Multiagent Systems· Joss Armstrong·· 4 hr agoAI score38
MARGIN:多智能体基础模型协同的运行时置信度校准
MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination
AI brief
论文提出 MARGIN(Multi-Agent Runtime Grading via Incremental Normalisation),一种运行时校准方法,可从观测到的答案结果中学习各模型特定的置信度修正,无需重新训练模型或预留校准集。
Source: arXiv · Multiagent Systems · arxiv.org