
Google DeepMind is testing a Gemini model with a new double-blind evaluation.
Cryptographic protection keeps test questions and model weights secret.
This could prevent benchmark contamination and support secure AI testing.
What happened
Google DeepMind is launching the first double-blind evaluation of a proprietary frontier AI model. For the pilot, it is testing a model from the Gemini Flash Lite line against confidential benchmarks, using Google Cloud's Confidential Space to keep both the test questions and the model private.
Why it matters
This method aims to solve benchmark contamination, where a model has already seen test questions during training, making results unreliable. It also removes the tradeoff where evaluators had to either share their test prompts or the provider had to share model weights. Google says this matters most for highly sensitive evaluations, such as cybersecurity or tests by government agencies.
What to watch
Google hopes this cryptographic approach sets a new standard for model oversight and helps build more reliable AI systems. The company details methodology and results in a technical report.
Ask the AI about this article →
The new double-blind evaluation from Google DeepMind addresses a long-standing problem in AI benchmarking. Previously, external evaluators had to choose between revealing their test questions or receiving the model's weights, which could compromise intellectual property. The recent delayed evaluation of Anthropic's Fable 5 for the ARC-AGI benchmark illustrates this dilemma, as the company's 30-day data retention policy for its strongest models created friction. By using Confidential Space from Google Cloud, the setup cryptographically verifies that both the test data and the model remain private, eliminating the need for such compromises.
This method also tackles benchmark contamination, where models that have seen test questions during training can produce inflated results. The cryptographic protection ensures that a model cannot optimize itself specifically for a test it has already seen. The pilot on a Gemini Flash Lite model against confidential benchmarks is a first step, but it could set a new standard for model oversight. If successful, this approach might allow independent organizations, especially in sensitive areas like cybersecurity or government, to rigorously test advanced models without sacrificing data sovereignty or security, as DeepMind hopes.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Home Depot is bringing its Magic Apron AI assistant to all U.S

Lam Research has started construction on an R&D center in Tualatin, Oregon, as part of a $3 billion, five-year…

ECRI, a nonprofit patient safety organization, is expanding its reporting system to include errors involving a…

Dozens of protesters gathered Thursday outside 8VC, a venture capital firm in Austin, to protest Palantir's AI…

Visa says its AI 'harness' makes Anthropic cheaper to use for cyber defense

AMD introduced ROCm 10, the next version of its software stack for AMD Instinct GPUs, designed to streamline t…
