arXiv · Software Engineering· Thiago Santos de Moura, Fynn Matuschek, Flavio Toffalini, Yannic Noller·· 3 小时前AI 评分50
LLM 生成代码的功能与安全差距纵向研究
Newer and Bigger, but Safer? A Longitudinal Study of the Functionality-Security Gap in LLM-Generated Code
AI 导读
论文对 32 个 LLM 做纵向研究,覆盖七个模型家族、三代旗舰与紧凑变体,用 CWEval 的 119 个任务、五种语言和 31 个 CWE 比较功能与安全差距。
来源:arXiv · Software Engineering · arxiv.org