用SAE与线性探针防御VLA对抗补丁
Get the story
该研究提出针对视觉-语言-动作(VLA)模型对抗补丁的机制化防御方法Detect and Suppress。研究使用稀疏自编码器(SAE)对VLA模型的表征进行机制化分析,识别出与对抗补丁存在强相关的特征,并仅在线性探针检测到攻击时,于推理阶段抑制该特征。该方法无需微调VLA模型。在LIBERO-10上的评估显示,间歇性干预可提升攻击下的成功率,而持续施加同一干预会显著降低策略性能。结果表明,攻击相关内部表征可作为VLA对抗防御的靶点,干预时机对避免干扰正常策略行为至关重要。
Generated from reports · updated 16 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · RoboticsDetect and Suppress:针对 VLA 模型对抗补丁的机制化防御
研究用稀疏自编码器(SAE)机制化分析 VLA 表征,识别出与对抗补丁存在强相关的特征,并仅在线性探针检测到攻击时于推理阶段抑制该特征。该方法无需微调 VLA,在 LIBERO-10 上评估显示间歇性干预可提升攻击下成功率,而持续施加同一干预会显著降低策略性能。结果表明攻击相关内部表征可作为 VLA 对抗防御靶点,干预时机对避免干扰正常策略行为至关重要。
Heat trend
Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.