Hot eventLive
研究评估13个LLM模拟新手程序员误解能力
1 reports1 sources3 hr ago updated
Get the story
AI overview
2026年10月9日,arXiv Software Engineering 频道发表一项研究,评估13个大语言模型在代码生成、求解、模拟与诊断新手编程误区上的表现,聚焦代码追踪问题。研究发现,前沿模型能可靠完成这些任务,但参数规模不超过14B的小模型在动态执行上表现吃力;经过代码微调的模型在模拟误区时容易退回正确执行。此外,误区诊断在True/False格式下显著比开放式生成更容易。该研究为目前唯一报道,暂无更早或相互矛盾的信息。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
Oct 9, 2026
- arXiv · Software EngineeringLLM 能否模拟新手程序员的认知误区?
研究评估了 13 个 LLM 在代码生成、求解、模拟与诊断新手编程误区上的表现,聚焦代码追踪问题。前沿模型能可靠完成这些任务,但 ≤14B 的小模型在动态执行上吃力,代码微调模型在模拟误区时会退回正确执行。误区诊断在 True/False 格式下显著比开放式生成更容易。
Heat trend
Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.