Hot eventLive
MatrixReward基于评判矩阵构建开放式生成奖励
1 reports1 sources3 hr ago updated
Get the story
AI overview
MatrixReward 通过逐条 rubric 的 pairwise 比较构建 rollout-by-rubric win-rate 矩阵,利用列间区分度和相关性计算数据依赖的 rubric 权重,并结合先验权重经列归一化后定义正负理想轮廓,以相对接近度作为质量奖励。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
Oct 2, 2026
- arXiv · Statistics Machine LearningMatrixReward:基于评分矩阵构建开放式生成奖励机制
MatrixReward 通过逐条 rubric 的 pairwise 比较构建 rollout-by-rubric win-rate 矩阵,利用列间区分度和相关性计算数据依赖的 rubric 权重,并结合先验权重经列归一化后定义正负理想轮廓,以相对接近度作为质量奖励。
Heat trend
There is not enough continuous observation data to draw a trend yet.