Google DeepMind tests double-blind evaluation of Gemini Flash Lite to counter benchmark contamination
Google DeepMind is piloting the first double-blind evaluation of the proprietary model Gemini Flash Lite in a cryptographically secured environment with partners such as Singapore AI Safety Institute, OpenMined, AVERI and MLCommons to prevent benchmark contamination.
Companies and institutions that rely on benchmark results when selecting or auditing AI models now have a precedent for a methodology intended to reduce the risk of distorted scores caused by contamination of test questions.