PowerCodeBench:电力系统代码生成知识边界探测基准
Get the story
2026年10月8日,arXiv 软件工程方向发布一手研究,提出面向 LLM 电力系统代码生成的知识边界探测与需求引导干预方法。研究推出 PowerCodeBench,这是一个面向 pandapower 的 2,000 任务参数化基准,用于探测模型在电力系统代码生成中的知识边界。配套提出免权重更新的部署期工作流:通过文档驱动的 L0-L3 探测生成逐模型 API 画像,在生成前按需求注入分层 API 证据,并在执行后进行定向修复。该工作不涉及模型权重更新,强调在部署阶段利用文档与执行反馈改善代码生成质量。目前报道仅披露上述方法与基准规模,未给出实验结果细节。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Software Engineering面向 LLM 电力系统代码生成的知识边界探测与需求引导干预
研究提出 PowerCodeBench(面向 pandapower 的 2,000 任务参数化基准)与免权重更新的部署期工作流,用文档驱动的 L0-L3 探测生成逐模型 API 画像,并在生成前按需求注入分层 API 证据、执行后定向修复。
Heat trend
Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.