AIToday

Microsoft's AI security tools score 12 points higher than Anthropic, Google, OpenAI on benchmark

Ars Technica AI1h agoSend on LINE
Microsoft's AI security tools score 12 points higher than Anthropic, Google, OpenAI on benchmark

Key takeaway

Microsoft announced two new AI-powered security tools on Monday: MDASH with MAI-Cyber-1-Flash, which scored 96 percent on the CyberGYM benchmark—12 points higher than competing systems from Anthropic, Google, and OpenAI—and costs half as much as the previous version; and Project Perception, a platform of specialized AI agents designed to handle 90 percent of security tasks at lower cost than competitors. The tools address what Microsoft describes as a fundamental shift in cybersecurity, where AI is accelerating both attack speed and scale, forcing defenders to protect increasingly complex networks with conventional methods.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Microsoft announced two AI security tools on Monday—MDASH with MAI-Cyber-1-Flash, which scored 96 percent on the CyberGYM benchmark (12 points higher than Anthropic's Mythos, and ahead of Google Gemini and OpenAI GPT), and Project Perception, a platform of specialized AI agents for red-, blue-, and green-team security functions. The new MDASH costs half as much to use as the previous MDASH offering.

  • Why it matters

    Microsoft says Project Perception can perform 90 percent of security tasks for lower costs than competing platforms, leaving organizations to use costlier alternatives only for the remaining 10 percent of tasks. The tools respond to what Microsoft calls a seismic shift in how organizations defend against cyberattacks, as AI accelerates both the speed and scale of threats, forcing security teams to manage increasingly complex environments with outdated approaches.

  • What to watch

    Both tools are currently in preview mode and should be closely scrutinized before use in production environments, though there are clear risks to not adopting such tools either. The balance between the risks of using AI agents versus the threat of avoiding them remains unresolved.

In Depth

Microsoft announced two new AI-powered cybersecurity tools designed to help organizations defend against rapidly accelerating threats. The first, MDASH with MAI-Cyber-1-Flash, demonstrated strong performance on CyberGYM, a standard benchmark test, earning a 96 percent score. This result places the tool 12 points ahead of Anthropic's Mythos and also outperforms offerings from Google (Gemini) and OpenAI (GPT). In addition to its performance advantage, the new MDASH version costs half as much to deploy as the previous MDASH offering, making it a more economical choice for cost-conscious organizations.

The second tool, Project Perception, takes a different approach by combining multiple specialized AI agents that handle distinct security functions. These agents perform red-team operations (simulating attacks to find vulnerabilities), blue-team operations (investigating vulnerabilities to assess risk), and green-team operations (implementing corrective actions). Rather than relying on a single model for all tasks, Project Perception automatically selects the most appropriate model based on the work at hand. Microsoft said this selection process factors in both the model's effectiveness and the cost to the customer, with decisions shaped by "ongoing research, benchmarking and evaluation across frontier and specialized models."

Microsoft framed these tools as a response to fundamental changes in the security landscape. The company stated that "as AI accelerates the speed and scale of cyberattacks, defenders are being asked to secure increasingly complex digital environments with approaches built for a different era." According to the announcement, security teams often struggle to piece together signals, context, and risk insights across vast amounts of data, making it harder to keep pace with emerging threats. Project Perception aims to address this challenge by performing 90 percent of security tasks at lower costs than competing platforms, with organizations only needing to resort to more expensive alternatives for the remaining 10 percent of use cases.

Both tools are currently in preview mode. The article cautions that they deserve careful scrutiny and evaluation before deployment in production environments, noting both the clear risks of adopting AI agents in security roles and the equally significant risks of avoiding such tools altogether. The balance between these competing risks remains unresolved.

Context & Analysis

Microsoft's announcement reflects an industry-wide tension: as artificial intelligence accelerates both the sophistication and speed of cyberattacks, traditional security approaches have become insufficient. The company frames the problem as one of scale and complexity—security teams must now process vast amounts of data and correlate signals across increasingly complex digital environments, yet they lack tools designed for this AI-driven threat landscape. By positioning MDASH and Project Perception as cost-effective alternatives to existing platforms, Microsoft suggests that price and performance have become differentiators in the security AI market.

The benchmark comparison—a 12-point advantage over Anthropic's Mythos and leads against Google and OpenAI's flagship models—serves as a key competitive claim, though the article notes that these tools warrant careful evaluation before production deployment. The emphasis on cost efficiency in Project Perception (performing 90 percent of tasks at lower expense) also points to a broader market shift: buyers are no longer purchasing monolithic security solutions but instead seeking platforms that intelligently allocate expensive, specialized models only where they are most needed. Both tools are in preview, indicating that real-world validation is still underway.

FAQ

How does the new MDASH compare to competitors on the CyberGYM benchmark?
MDASH with MAI-Cyber-1-Flash scored 96 percent on CyberGYM, which is 12 points higher than Anthropic's Mythos, and also beats Google Gemini and OpenAI GPT on the same test.
What does Project Perception do?
Project Perception is a collection of specialized AI agents that perform red-team (attack simulation), blue-team (defense), and green-team (corrective action) functions for finding vulnerabilities, investigating their risk, and taking corrective actions respectively. The platform automatically selects which models to use based on the assigned task, considering both model effectiveness and customer cost.
What cost advantage does Project Perception claim?
Microsoft said Project Perception can perform 90 percent of security tasks for lower costs than similar platforms from competitors, leaving customers to use more expensive alternatives only for the remaining 10 percent of tasks.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime