arXiv · Computation and Language· Hala Almaghout, Christian Federmann, Qin Gao·· 4 hr agoAI score38
大语言模型用于机器翻译质量标注:人类与模型均面临挑战
Large Language Models for Machine Translation Quality Annotation: Humans and Models Are Both Challenged
AI brief
论文评估 LLM 在 MQM 与 ESA 两种机器翻译质量评估方案中的表现,在 70 个语言对的长上下文测试集及 WMT23、WMT25 数据上比较其与人类标注者的一致性。
Source: arXiv · Computation and Language · arxiv.org