Google pilots double-blind AI evals to stop benchmark cheating
Two moves this hour quietly attack the same weakness in how we trust AI: how to know a model is actually good — and how platform ages are confirmed. Google thinks cryptography can fix the first; Meta is leaning on Apple and Google to fix the second.
Google DeepMind is piloting the world's first double-blind AI evaluations, keeping external test prompts sealed in a cryptographically protected "box" so a frontier model can't peek at questions before it sits the exam. The mechanism targets benchmark contamination — when a model has already seen a test's prompts, its score gets artificially inflated and no longer reflects real capability. Today that forces an ugly tradeoff on outside evaluators: hand over your test questions (and risk the model provider absorbing them into training) or hand over your weights (and expose proprietary work). DeepMind's answer lets independent labs, civil-society groups and national AI Safety and Security Institutes test models rigorously with neither side surrendering its secrets, enforced by technical and cryptographic safeguards rather than contracts alone. The lab sees this as a major step precisely because high-stakes evaluations — cybersecurity checks, government review — are where contamination does the most damage. It's a low-drama technical fix, but it's quietly one of the more meaningful credibility improvements in benchmarking in years: the industry has spent months arguing over who can be trusted to grade the models, and this gives outside auditors a way to do it that doesn't demand existential trust.
Meta's $18 billion child-safety settlement will verify age using "reliable age signals" shared by Apple and Google's operating systems and app stores. Buried in the deal Meta struck with state attorneys general on the 26th is a structural bet: instead of building its own age checks from scratch, Meta's framework will lean on signals Apple and Google surface from their OS and store layers — a recognition, per the settlement text, that the platforms controlling the devices are best placed to confirm who's on them. It lands against wider skepticism that age verification actually works well yet: biometric, ID and behavioral methods each carry flaws, and handing sensitive identity data to third parties raises its own privacy risks. The design points toward token-based checks that confirm "under 13" or "under 16" while discarding the personal details — a model privacy advocates broadly prefer over storing scans of faces. The settlement's real test is whether the immature tech behind it can carry the most detailed youth-safety rules a major platform has ever accepted, especially with a federal trial still looming over whether Meta's platforms hooked minors in the first place.
What to watch: whether other platforms adopt Google's double-blind evaluation standard, and which OS-level age signals Apple and Google actually expose to Meta.
Do you trust cryptographic third-party evals more than a lab grading its own models — and should age verification live at the OS level? Tell us in the comments.
Sources: Google DeepMind · Techmeme · TechCrunch · The Verge · Techmeme page 35