Skip to content
Hot eventLive

MoE路由logits用于VLM生成前安全检测

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026年10月7日,arXiv Computation and Language 发表一手研究,提出一种轻量级 router-logit 安全检测器。该方法在 prompt prefill 阶段读取 MoE 视觉语言模型的路由信号,在生成前识别不安全请求,不修改模型参数或专家路由。研究称该检测器在 Qwen3-VL 和 Kimi-VL 上显著降低 HoliSafe 基准的安全错误,并泛化至 MISHard 和 MM-SafetyBench 等分布外安全基准。目前未见与早先报道矛盾之处,也无后续独立验证报道。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 7, 2026
  1. arXiv · Computation and Language
    阅读而非操控:利用 MoE 视觉语言模型中的 Router Logits 提升多模态安全性

    研究提出一种轻量级 router-logit 安全检测器,在 prompt prefill 阶段读取路由信号,在生成前识别不安全请求,不修改模型参数或专家路由。该检测器在 Qwen3-VL 和 Kimi-VL 上显著降低 HoliSafe 基准的安全错误,并泛化至 MISHard 和 MM-SafetyBench 等分布外安全基准。

Heat trend

Current heat 9·Comparable peak 10(Oct 7)·Comparable change over 24 hours –

02.557.510Oct7Oct7Oct7Oct7

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.