Skip to content
Hot eventLive

Kimi K2.7 Code经RL训练后在多个编码基准上显著提升

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026-10-02 arXiv发布软件工程领域研究:研究者对Kimi K2.7 Code(1T参数,32B激活)进行单epoch GSPO RL后训练,仅用1700个任务,即在SWE-Bench Pro、DeepSWE、Terminal-Bench 2.1/3/4和SWE-Marathon六个外部基准上提升pass@1,且对训练后发布的三个数据集仍显著(p=0.004)。

Generated from reports · updated 2 hr ago

Timeline

Follow the coverage from different angles.

Oct 2, 2026
  1. arXiv · Software Engineering精选
    基于智能体编码任务的 RL 跨基准迁移研究

    研究者对 Kimi K2.7 Code(1T 参数,32B 激活)进行单 epoch GSPO RL 后训练,仅用 1700 个任务,即在 SWE-Bench Pro、DeepSWE、Terminal-Bench 2.1/3/4 和 SWE-Marathon 六个外部基准上提升 pass@1,且对训练后发布的三个数据集仍显著(p=0.004)。

Heat trend

There is not enough continuous observation data to draw a trend yet.