组合偏好与规则奖励后训练文生图模型
Get the story
2026年10月5日,arXiv Computer Vision(一手)报道,研究者提出一种面向开放域文生图的后训练方案,将偏好奖励与基于准则的奖励组合,以兼顾人类审美偏好、提示词忠实性并抑制奖励破解。RL训练后的Flux2dev在Arena文生图榜上Elo较基础模型高69分;后训练的Ideogram-4达到Elo 1223.5,超过榜单所有开源模型(截至2026年9月4日快照)。目前未见后续报道或独立验证。
Generated from reports · updated 16 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Computer Vision通过组合偏好与准则奖励后训练前沿文生图模型
研究者提出一种面向开放域文生图的后训练方案,将偏好奖励与基于准则的奖励组合,以兼顾人类审美偏好、提示词忠实性并抑制奖励破解。RL 训练后的 Flux2dev 在 Arena 文生图榜上 Elo 较基础模型高 69 分,后训练的 Ideogram-4 达到 Elo 1223.5,超过榜单所有开源模型(截至 2026 年 9 月 4 日快照)。
Heat trend
Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.