SAFESHIELD:小模型部署期安全决策组织框架
Get the story
2026年10月7日,arXiv Software Engineering 发布一手研究,提出 SAFESHIELD 框架,将小型语言模型的部署期安全建模为决策组织问题。该框架把安全决策拆分为准入、路由、证据、发布四类职责,并以可审计的 Decision Traces 记录已提交决策。协调消融实验显示,切断准入门控会显著增加误发布;当发布决策缺少上游证据时,发布准确率从 96.0% 降至 69.5%。研究结果表明,部署期安全不仅取决于单个护栏能力,也取决于决策如何组织与协调。目前报道仅涉及该框架设计与实验结果,未见后续验证或部署信息。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Software EngineeringSAFESHIELD:面向小型语言模型部署期安全的决策组织框架
SAFESHIELD 把小型语言模型的部署期安全建模为决策组织问题,将安全决策拆分为准入、路由、证据、发布四类职责,并以可审计的 Decision Traces 记录已提交决策。协调消融实验显示,切断准入门控会显著增加误发布;发布决策缺少上游证据时,发布准确率从 96.0% 降至 69.5%。结果表明部署期安全不仅取决于单个护栏能力,也取决于决策如何组织与协调。
Heat trend
Current heat 9·Comparable peak 10(Oct 7)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.