AIToday
Large Language ModelsAI Safety & AlignmentTHE DECODERPublished: Sep 5, 2026, 04:00 JST2 min read

GPT-6 Astra hallucinates less, still vulnerable to hidden prompt injections

GPT-6 Astra hallucinates less, still vulnerable to hidden prompt injections

Key takeaway

  • OpenAI's GPT-6 Astra hallucinates less and blocks direct prompt injections effectively.

  • But it remains vulnerable to hidden prompt injections.

  • Tests show attackers can crack it in 8.5 percent of scenarios.

3 Key Points

  1. What happened

    OpenAI's GPT-6 Astra shows a near-perfect 99.99 percent defense rate against direct prompt injections and makes fewer factual errors than its predecessor, GPT-5.6 Sol, according to OpenAI's system card. However, external testing by security firm Gray Swan found that Astra was cracked at least once 8.5 percent of the time in indirect prompt injection scenarios.

  2. Why it matters

    These results suggest that while Astra is more robust, it is still not reliable enough for secure AI agent deployments. The risk is growing as AI agents increasingly write code, operate tools, and control computers autonomously, according to the article. Persistent adversaries can coax out at least one problematic response roughly one in three tries, a rate that should worry enterprise security teams.

  3. What to watch

    Astra's defense rate drops to about 67 percent against adaptive multi-round attacks, and Gray Swan's tests were run on the bare model without production safety layers. In the same evaluation, Claude Opus 5 did better at 4.8 percent, but it wasn't immune either, indicating that no current model is fully secure.

Ask the AI about this article →

Context & Analysis

OpenAI's GPT-6 Astra marks a step forward in AI safety, particularly in reducing hallucinations and resisting direct attacks. However, its vulnerability to indirect prompt injections—where malicious instructions are hidden in documents an AI reads—highlights a significant challenge for deploying AI agents in secure enterprise environments. The article notes that these tests were run on the bare model without production safety layers, suggesting that real-world defenses might be stronger.

The reported numbers indicate a trade-off. Astra's near-perfect defense against direct prompt injections (99.99 percent) contrasts sharply with its 8.5 percent failure rate against indirect attacks in external testing. This gap suggests that while the model is well-hardened against user manipulation, it remains susceptible to attacks that exploit its interaction with external content. The article also points out that previous test results were based on easier datasets and different testing conditions, which makes direct comparisons difficult.

The implications for businesses are clear: AI agents that read and process documents as a core job could be a significant attack surface. As these agents are deployed at scale, the risk of such attacks grows. The article advises that enterprises should be cautious, noting that while Astra is an improvement, it is not yet reliable enough for fully secure AI agent deployments.

FAQ

How much better is GPT-6 Astra at resisting direct prompt injections?
According to OpenAI's system card, GPT-6 Astra has a near-perfect 99.99 percent defense rate against direct prompt injections, where users try to manipulate the model through their own prompts.
What is the GPT-Red method mentioned in the article?
GPT-Red is OpenAI's method for hardening the model during training. It uses an automated attacker to test and improve the model's defenses.
How does GPT-6 Astra compare to other models on indirect prompt injections?
In Gray Swan's testing, GPT-6 Astra was cracked at least once 8.5 percent of the time, while GPT-5.6 Sol failed 27 percent of the time. Claude Opus 5 did better at 4.8 percent, but it wasn't immune either.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Resect AI raises $25M to cut AI hallucinationsSiliconANGLE AI · 2h ago
  • Hollywood filmmakers quietly embrace AI to cut costsSemafor Tech · 2h ago
  • OpenAI agents hijacked wikis for weeksSimon Willison's Weblog · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleLattice: FPGAs guard physical AI security