Skip to content
Hot eventLive

SSRFT:将安全对齐重构为安全角色内化

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026年10月7日,arXiv(Artificial Intelligence,一手来源)报道了SSRFT(Supervised Safe-Role Fine-Tuning)方法,提出将大语言模型的安全对齐重构为对预定义安全角色的内化,而不是依赖显式的拒绝模式。该方法使用心理测量问题、少量越狱提示与安全角色描述构建SRQA数据集,并在此基础上合成并扩展与角色一致的回复,以训练模型在安全角色框架内作出回应。目前公开信息仅涉及该方法的基本思路与数据构建方式,尚无后续实验结果或第三方验证报道。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 7, 2026
  1. arXiv · Artificial Intelligence
    SSRFT:把安全对齐重构为安全角色内化,超越拒绝模式

    SSRFT(Supervised Safe-Role Fine-Tuning)把 LLM 安全对齐重构为预定义安全角色的内化,而非依赖显式拒绝模式。它用心理测量问题、少量 jailbreak 提示与安全角色描述构建 SRQA 数据集,合成并扩展角色一致回复。

Heat trend

Current heat 9·Comparable peak 10(Oct 7)·Comparable change over 24 hours –

02.557.510Oct7Oct7Oct7Oct7

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.