Adapted from clinical-trial design, double-blind AI evaluation keeps model weights and inference code hidden from the evaluator while keeping test prompts and scoring logic hidden from the model provider. Google DeepMind's August 2026 pilot used Google Cloud Confidential Space, an NVIDIA H100 secure enclave, and OpenMined PySyft with partners including Singapore AISI, AVERI, and MLCommons. Only aggregate metrics exit the enclave after both parties approve redacted code.