Google DeepMind tests double-blind evaluation of Gemini Flash Lite to counter benchmark contamination
Google DeepMind is piloting the first double-blind evaluation of the proprietary model Gemini Flash Lite in a cryptographically secured environment with partners such as Singapore AI Safety Institute, OpenMined, AVERI and MLCommons to prevent benchmark contamination.
According to its statement, Google DeepMind has launched a pilot program for the first double-blind evaluation of a proprietary frontier-class model – specifically Gemini Flash Lite. The test questions are sealed in a cryptographically secured environment (a “box”) that the model cannot access in advance, so it cannot use them to optimize its results before testing takes place.
The aim is to address the problem of so-called benchmark contamination – a situation in which the model has already “seen” the test questions and the resulting score does not reflect its actual capabilities. Google DeepMind is collaborating on the project with five external partners: Singapore AI Safety Institute, OpenMined, AVERI and MLCommons. Testing takes the form of confidential benchmarks in an environment designed to protect the privacy of the test data.
According to Google DeepMind, it has so far relied on protocols without logging and contractual confidentiality guarantees to protect test prompts. The company describes the introduction of technical and cryptographic security measures as a significant step forward in securing model evaluation, because it goes beyond contractual guarantees alone.
Why it matters
If a model can see test questions in advance, its benchmark scores overstate its actual capabilities and mislead those who rely on the results – policymakers, researchers and companies selecting AI providers. According to Google DeepMind, cryptographically securing test questions instead of relying solely on contractual guarantees is intended to increase the credibility of evaluations and potentially serve as a template for evaluating other proprietary models.
Relevant practical impact
What this means
For a business
Companies and institutions that rely on benchmark results when selecting or auditing AI models now have a precedent for a methodology intended to reduce the risk of distorted scores caused by contamination of test questions.
Risks and complianceCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.