arXiv · Software Engineering· Thiago Santos de Moura, Fynn Matuschek, Flavio Toffalini, Yannic Noller·· 4 hr agoAI score50
LLM 生成代码的功能与安全差距纵向研究
Newer and Bigger, but Safer? A Longitudinal Study of the Functionality-Security Gap in LLM-Generated Code
AI brief
论文对 32 个 LLM 做纵向研究,覆盖七个模型家族、三代旗舰与紧凑变体,用 CWEval 的 119 个任务、五种语言和 31 个 CWE 比较功能与安全差距。
Source: arXiv · Software Engineering · arxiv.org