DeepMind announced a new initiative to conduct double-blind evaluations of its advanced AI models, a method designed to improve the integrity and trustworthiness of AI performance benchmarks. This approach seeks to address the problem of "benchmark contamination," where AI models may inadvertently "peek" at test data, leading to artificially inflated scores that do not accurately reflect their true abilities. The company is collaborating with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons to implement these evaluations.

The core of DeepMind's new evaluation system involves confining external evaluations within a cryptographic "box." This secure environment prevents AI models from accessing or optimizing their performance based on test questions prior to formal evaluation. DeepMind emphasized that while it conducts extensive internal testing, external evaluations are crucial for identifying potential blind spots in its AI systems. These external partners bring specialized expertise to stress-test models.

The company highlighted the increasing importance of preventing benchmark contamination as AI models become more capable. Policymakers, researchers, and enterprises rely on AI benchmarks to accurately gauge model capabilities and safety. If models can preview evaluation questions, the resulting scores may not be trustworthy. DeepMind previously utilized rigorous contractual safeguards and zero-logging protocols to maintain the confidentiality of external test prompts.

This double-blind methodology is analogous to high-stakes examinations where students must not see test questions in advance for their scores to be meaningful. DeepMind is applying this principle to AI, specifically with a Gemini Flash Lite model, which will be tested against confidential benchmarks in a privacy-preserving setting. This move aims to increase the overall integrity of AI evaluation processes.

DeepMind has a history of conducting rigorous evaluations for its AI systems. For example, in May 2026, the company published benchmark results for its AI Co-clinician, which involved blind head-to-head evaluations with physicians. In that instance, 67% of physicians preferred the AI Co-clinician over existing clinical tools when they were unaware of which system was which. This earlier work demonstrates a commitment to blind evaluation methods in specific applications.

The company also recently released findings and an evaluation framework to measure AI's potential for harmful manipulation in areas like finance and health. This research, published in March 2026, involved nine studies with over 10,000 participants and focused on understanding if AI could influence decision-making in simulated scenarios. DeepMind also introduced an exploratory Harmful Manipulation Critical Capability Level within its Frontier Safety Framework to track models with capabilities that could systematically alter beliefs and behaviors. This broader effort to ensure AI safety and reliability underpins the current move towards double-blind evaluations.

The deployment of double-blind evaluations represents an attempt to create a more standardized and verifiable method for assessing advanced AI. By preventing models from having any prior exposure to the evaluation criteria, DeepMind seeks to provide a clearer picture of their inherent capabilities, rather than their ability to adapt to known tests. This development could influence how other AI developers approach model assessment, potentially leading to more transparent and trustworthy benchmarks across the industry.