跳到正文
原文
arXiv · Information Retrieval· Omar Sheta, Rinku Deuja, Hadi Masoudi, Minghong Fang·· 6 小时前精选AI 评分60

不安全梯度能否在对话中存活?基于梯度的越狱检测在多轮对话中的脆弱性

Does the Unsafe Gradient Survive a Conversation? On the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue

AI 导读

论文评估了单提示梯度越狱检测器 GradSafe 在多轮对话中的表现,将其扩展为 Context Window Scanner,用固定大小窗口的最大得分作为对话级分数。

推荐理由

梯度检测在合成数据上表现稳健,但在真实良性对话中大幅退化,提示实际部署需重新校准。

来源:arXiv · Information Retrieval · arxiv.org