AIToday
arXiv cs.RO (Robotics)Published: Apr 27, 2026, 13:00 JST1 min read

RedVLA: First red teaming framework for detecting unsafe behaviors in Vision-Language-Action models before deployment

3 Key Points

  1. Researchers propose RedVLA, a two-stage red teaming framework designed to systematically uncover unsafe physical behaviors in VLA models (AI systems that combine vision, language, and robotic action capabilities). The framework uses Risk Scenario Synthesis to construct initial risk scenes by identifying critical interaction regions from benign trajectories, then applies Risk Amplification through gradient-free optimization to ensure stable unsafe behavior elicitation.

  2. Experiments on six representative VLA models show RedVLA achieves an Attack Success Rate (ASR) up to 95.5% within 10 optimization iterations, uncovering diverse unsafe behaviors. The researchers also propose SimpleVLA-Guard, a lightweight safety guard built from RedVLA-generated data to mitigate these risks.

  3. Code, data, and assets from the study are made available publicly, enabling further development of safety mechanisms for VLA model deployment.

Ask the AI about this article →

arXiv cs.RO (Robotics)Read Original Article

Get AI news like this every morning

For example, today's edition would include:

  • ZeroDrift launches Guard for Agents compliance serviceSiliconANGLE AI · 2h ago
  • Zhen Ding unveils 1.6T optical modules, 34-layer AI boardsDIGITIMES Asia · 2h ago
  • Semicon Taiwan 2026 to test AI spending boomDIGITIMES Asia · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Next articleResearchers introduce V-STC, a multi-vehicle coordination method that reduces temporal occupancy while maintaining collision-free separation