Hot eventLive
论文研究分类后缀提升激活探针的恶意输入检测泛化
1 reports1 sources16 hr ago updated
Get the story
AI overview
该研究探讨在用户回合后追加分类指令,以提升激活探针对越狱与提示注入的分布外检测。报道显示,单位置探针最高约提升 4 个 AUC 点,且收益来自分类格式而非具体类别名;结论在 13 个安全基准、Llama-3.1-8B、Qwen3.5-9B、Gemma-4-12B 的留一数据集评估中成立,并延续到 attention、multi-max、MLP 等多位置池化探针。
Generated from reports · updated 15 hr ago
Timeline
Follow the coverage from different angles.
Oct 5, 2026
- arXiv · Machine Learning Theory提示词如何提升激活探针在野外的恶意输入泛化检测
在用户回合后追加分类指令可让激活探针提升越狱与提示注入的分布外检测,单位置探针最高约 4 个 AUC 点,且收益来自分类格式而非具体类别名。该结论在 13 个安全基准、Llama-3.1-8B、Qwen3.5-9B、Gemma-4-12B 的留一数据集评估中成立,并延续到 attention、multi-max、MLP 等多位置池化探针。
Heat trend
Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.