跳到正文
原文
arXiv · Machine Learning Theory· Gowthamkumar Nandakishore·· 3 小时前AI 评分37

CLM-as-a-Judge:在公开 Judge 基准上评测开源对比决策模型

CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks

AI 导读

开源对比决策模型 Contrastive-LM/CLM-v0.1-8B 在公开评测基准上作为评判接近随机:best-of-four 得分 0.351(随机 0.250)。

来源:arXiv · Machine Learning Theory · arxiv.org