Source · Tier B · Technical report

Double Blind Evals: Resolving the Dual Confidentiality Dilemma in AI Safety Auditing

A. Trask, S. Messing, V. Pahwa, P. Maham, R. Kolga, A. Frantz, A. Tash, K. Thomas, S. McGregor, G. Balston, P. Paskov, M. Brundage, A. Vij, B. Hillenbrand, A. Karargyris, T. Acosta, J. Fenster, M. Eilish, R. Elasmar, M. Khan, K. van der Veen, R. S, S. Wagh, S. Gabriel, P. Werneck, L. Strahm, K. McDonough, R. Falcon, K. Lum, W. Isaac. 2026. Google DeepMind.

Originalhttps://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/piloting-the-worlds-first-double-blind-ai-evaluations/double-blind-evaluations-technical-report.pdf
VersionTechnical report linked from Google DeepMind's blog post "Piloting the world's first double-blind AI evaluations" (https://deepmind.google/blog/piloting-the-worlds-first-double-blind-ai-evaluations/, dated 2026-08-27). The PDF prints no date. Author affiliations as printed: Google, AVERI, Singapore AISI, OpenMined and ML Commons. No arXiv version was found on 2026-09-25.
Accessed2026-09-25
Authored or published byOpenMined
NoteReport by the pilot's participants on their own system; supports "the participants report" statements. Gemini 2.5 Flash Lite was evaluated against private AILuminate prompts in a GCP Confidential Space enclave (a3-highgpu-1g, one NVIDIA H100 with Intel TDX), using OpenMined's PySyft; both parties verified the attestation before uploading assets.

Cited by

Search

Full search page