Hot eventLive
Kimi K2.7 Code经RL训练后在多个编码基准上显著提升
1 reports1 sources3 hr ago updated
Get the story
AI overview
2026-10-02 arXiv发布软件工程领域研究:研究者对Kimi K2.7 Code(1T参数,32B激活)进行单epoch GSPO RL后训练,仅用1700个任务,即在SWE-Bench Pro、DeepSWE、Terminal-Bench 2.1/3/4和SWE-Marathon六个外部基准上提升pass@1,且对训练后发布的三个数据集仍显著(p=0.004)。
Generated from reports · updated 2 hr ago
Timeline
Follow the coverage from different angles.
Oct 2, 2026
- arXiv · Software Engineering精选基于智能体编码任务的 RL 跨基准迁移研究
研究者对 Kimi K2.7 Code(1T 参数,32B 激活)进行单 epoch GSPO RL 后训练,仅用 1700 个任务,即在 SWE-Bench Pro、DeepSWE、Terminal-Bench 2.1/3/4 和 SWE-Marathon 六个外部基准上提升 pass@1,且对训练后发布的三个数据集仍显著(p=0.004)。
Heat trend
There is not enough continuous observation data to draw a trend yet.