Researchers evaluated 13 architectural configurations of multi-agent systems (networks of two or more autonomous AI agents) across browser, desktop, and code environments, using stagewise evaluations to measure planning refusal, execution-stage interception, partial harmful execution, and successful attack completion.
Multi-agent architectures were more vulnerable than standalone agents in the majority of configurations, with attack success rates varying by up to 3.8x at comparable or higher benign accuracy. Three key design choices shaped the security tradeoff: agent roles (how authority and responsibility are allocated), communication topology (how and when agents interact), and memory (context and state visibility accessible to each agent).
No single multi-agent design was found to be universally safer, indicating that architectural decisions governing agent coordination create attack surfaces that have not been systematically characterized until this empirical study.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic reset the 5-hour and 1-week usage limit windows for its AI service Claude on September 1, in connect…

Salesforce and Anthropic announced Claudeforce, starting with "Salesforce in Claude." This plugin lets users i…

Anthropic announced Claude Fable 5.1 and Claude Mythos 5.1 on September 1

A technical explainer compares three LLM serving strategies—static, dynamic, and continuous batching

Anthropic's latest model, Claude Fable 5.1, is now available on Snowflake Cortex AI

The Allen Institute for AI released BenchMIRT, a method to audit AI benchmarks question-by-question
