文本中心全模态后训练提升推理并降成本
Get the story
2026年10月5日,arXiv Computation and Language 发布论文,提出以文本为中心的全模态推理后训练范式:先用文本-only 监督微调与强化学习优化推理,再用减少数据的原生音视频 RL 精炼感知。论文称,文本-only SFT 加 RL 使 Qwen2.5-Omni-7B 在九项推理分数上的几何均值较基线提升 25.83%,比完整原生音视频路线少 56.6% GPU-hours;精炼阶段输入 token 减少约 90%,感知恢复至基线以上,并保留文本-only 管线 93.5% 的推理增益。
Generated from reports · updated 16 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Computation and Language论文提出以文本为中心的全模态推理后训练范式
论文提出以文本为中心的全模态推理后训练范式,用文本训练主优化推理,再用减少数据的原生音视频 RL 精炼感知。文本-only SFT 加 RL 使 Qwen2.5-Omni-7B 九项推理分数几何均值较基线提升 25.83%,比完整原生音视频路线少 56.6% GPU-hours;精炼阶段输入 token 减少约 90%,感知恢复至基线以上,并保留文本-only 管线 93.5% 的推理增益。
Heat trend
Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.