Hugging Face Blog·· Sep 3SelectedAI score69
用 GRPO 微调 350M 模型提升结构化输出表现
Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
AI brief
作者用 GRPO 和 TRL 对 LFM2.5-350M 进行约 100 步微调,在 IFStruct 基准上从 22.6% 提升到 29.7%。方案使用约 500 条样本、3 个奖励函数,并在免费 Colab 或 Kaggle GPU 上可运行,代码已开源。
Why it matters
原文给出了一套可在免费 GPU 上复现的小模型结构化输出微调方案,读者能直接迁移该奖励函数与数据增强思路。
Source: Hugging Face Blog · huggingface.co