Confidential multi-party verification

R2Demonstrated

Confidential multi-party verification runs an agreed check over assets that their owners will not show anyone, such as a developer's model weights, an auditor's test data or users' logs, so that no participant can see the others' inputs.

The check runs either inside hardware-isolated enclaves that sign what code ran or with zero-knowledge proofs. Only the result is released, with evidence of how it was produced. In a 2026 cloud GPU pilot, an outside evaluator tested a proprietary Gemini model on private prompts, neither side seeing the other's inputs.

Research prototypes compose multi-step audit workflows, limit usage monitoring to a jointly signed plan, and prove audits of small models with zero-knowledge proofs. No independent red-team of these systems has been published.

The main obstacles are scaling to frontier models, and trust in hardware vendors and the cloud host. Even a one-bit verdict can leak information about the private inputs.

Readinessmedium confidence

A 2026 pilot evaluated a proprietary model on a generally available cloud enclave service, but the evaluation workflows are pilots or research prototypes, and no party is documented as relying on their results for a verification decision.

Rubric assessment

Assessed use: audits or evaluations of a private model that reveal neither party's inputs

  • R1 met: designs that state what is verified and what is trusted are published for TEE workflows 1 2 and zero-knowledge audits 3.
  • R2 met. In the double-blind pilot, AVERI evaluated Gemini 2.5 Flash Lite against private benchmark prompts in a Google Cloud enclave with an NVIDIA H100, with the prompts kept from Google DeepMind and the weights from the evaluators (participant-reported) 9. Cove has an open-source reference implementation on Intel TDX via Phala Cloud's dstack, demonstrated end to end on a benchmark workflow 1 2. ZkAudit reports peer-reviewed end-to-end audits of MobileNet image classifiers and a recommender model 3. Attestable Audits ran safety benchmarks on an 8-billion-parameter model in AWS Nitro Enclaves 4.
  • R3 not met for this use. Google reports that Confidential Space, which releases each data owner's data only to an attested workload that meets the owner's conditions, is generally available, on H100 GPUs since April 2026 11 12. That release step is TEE attestation, assessed on its own record. The evaluation workflow that ran on it, OpenMined's PySyft, is documented only as a pilot 9. Pour Demain, an outside auditor, ran interpretability evaluations of a 744-billion-parameter open-weights model on Tinfoil's production confidential-computing platform, with its governance layer shown for single sessions only 10. No party is documented as relying on such a result for a verification decision.
  • R4 not met: no independent audit or red-team of Confidential Space for this use, Cove, Attestable Audits or ZkAudit has been published. An independent analysis of one monitoring tool's published evaluation found that its one-bit verdicts leak information 7. Attackers with physical access to the memory bus have forged the Intel TDX attestation the pilot relied on 13 14.

Confidence is medium. The R2 evidence includes peer-reviewed results 3 and a pilot on a proprietary model 9. Whether a general-purpose enclave service can carry this use to R3 is a judgment call.

Gaps to the next level
  • A production-grade, available confidential evaluation or audit workflow, or reliance by a party other than its developer on such a result for a verification decision.
  • An independent public evaluation (audit, red-team or peer-reviewed security analysis) of a confidential multi-party verification system, such as Confidential Space with PySyft or Cove, that leaves no critical flaw open.
  • Confidential evaluation on multi-node enclave clusters at the scale of the largest frontier models, or zero-knowledge audits at that scale.

Assessed 2026-09-28 against rubric v1.1.

On this page

How it works

Audits often need sensitive artifacts, such as model weights and evaluation data, held by parties who do not trust each other 1. Model providers have business reasons to keep models and data secret, while society needs algorithmic transparency 3. Confidential multi-party verification moves the agreed check to a place where no single party sees the others' inputs 1 3.

  • Composable enclave workflows. Cove uses trusted execution environments (TEEs) to compose reusable multi-stage audit workflows 1. In its developers' design, each data owner encrypts its artifact locally and approves only specific, hashed workflow steps 2. A key is released only to an enclave whose attestation matches an approved step 2. Each step emits a certificate that the hardware binds to its code and outputs, and a verifier checks the whole chain, starting from the final certificate 2.
  • Commercial confidential workspaces. Google reports that its Confidential Space runs a workload on data from several parties and releases each party's data only to a workload whose attestation meets that party's conditions. The operator has no access to the data 11. OpenMined's PySyft evaluation workflow used it in a 2026 pilot 9.
  • Enclave-run audits. In Attestable Audits, the model provider and the auditor each encrypt their inputs to an enclave key 4. The enclave runs the benchmark and publishes to a transparency log an attestation that binds the model hash, the hash of the audit code and data, and the result 4.
  • Plan-scoped monitoring. A monitoring party and a monitored party co-sign a plan, which an open-source classifier runs inside an attested TEE; scope changes need fresh signatures from both parties 5. The plan fixes the inputs, the computation steps and the outputs to be released 6.
  • Zero-knowledge audits. In ZkAudit, a provider publishes cryptographic commitments to its dataset and weights, plus a zero-knowledge proof that the weights came from training on that data 3. It then answers audit requests by computing a function privately and releasing the output with a proof that the computation was correct 3.
  • Designing what is disclosed. Minimal Information Disclosure treats the content of the released evidence as a design choice, and aims to minimize what the evidence reveals beyond the authorized result 7.

What it establishes

It can establish:

  • that an agreed computation ran on committed or attested inputs and produced the released result, without revealing the inputs 2 3;
  • in Attestable Audits, that the model answering a user is the one that was audited 4.

It cannot establish:

  • that nothing was left out of the workflow, since auditors also need assurance that all activity is accounted for 8;
  • that the released result reveals nothing else: in one analysis, a one-bit verdict let a predictor tell which of two groups a private request came from with 72.4% balanced accuracy, against 50% by chance 7;
  • in ZkAudit, secrecy of the model architecture, which the proofs reveal 3.

Threat model

  • Double-blind pilot. Its guarantees hold for a user who trusts that the cloud provider and the hardware manufacturer have insufficient incentives to collude against them 9. Google's services sign and verify the attestation report, which the participants note places Google in the verification path 9.
  • Cove. Its developers state that it trusts Intel TDX and its attestation roots, Docker's enforcement of the generated configuration, pinned first-party components, and public source code with human review 2. It does not trust its own storage server, network routing or orchestration 2.
  • Attestable Audits. All parties must trust the vendor of the confidential-computing technology, which was AWS in the prototype 4.
  • ZkAudit. It rests on cryptographic assumptions, including binding and hiding commitments 3.
  • Treaty setting. A governance analysis contrasts the confidential-computing adversary, a dishonest operator, with the adversary in a treaty setting: a state with physical access to data centres and leverage over hardware vendors 8.

Evidence

  • Double-blind pilot. Google DeepMind, OpenMined, AVERI, MLCommons and the Singapore AI Safety Institute report evaluating Gemini 2.5 Flash Lite against private AILuminate benchmark prompts in Google Cloud's Confidential Space, on one NVIDIA H100 with Intel TDX, using OpenMined's PySyft 9. Each party checked the enclave's attestation before uploading its assets, the prompts stayed hidden from Google DeepMind and the weights from the evaluators, and AVERI staff decrypted and scored the outputs 9. The participants note that not all model code could be inspected or allowlisted, and that the guest operating system builds are not independently reproducible 9. The report's figure of under 5% compute overhead is cited from an NVIDIA blog, not measured in the pilot 9.
  • Pour Demain on Tinfoil. Pour Demain reports running five interpretability workloads on GLM-5.1, a 744-billion-parameter open-weights model, on Tinfoil's production confidential-computing platform, using Intel TDX with eight H200 GPUs 10. Raw tensors stayed inside the enclave, and only bounded, aggregated, signed and budget-capped exports left it 10. It measured 5–7% added wall time from confidential computing, rising to 33–38% with interpretability instrumentation, and it demonstrated the governance layer for single sessions only 10.
  • Cove. Its authors show how its primitives express three applications: capability-attested inference, attested confidential benchmarks and bilateral capability verification 1. They report an open-source reference implementation on Intel TDX via Phala Cloud's dstack, demonstrated end to end on the benchmark workflow only 1 2.
  • Attestable Audits. The authors ran MMLU, XSum and ToxicChat on a 4-bit Llama-3.1-8B model in CPU-only AWS Nitro Enclaves 4. They report that CPU inference cost 21.7 times as much per token as GPU inference and ran about 100 times slower 4.
  • ZkAudit. Peer-reviewed at ICML 2024 3. The authors audited MobileNet v2 image classifiers on three datasets, with accuracy 0.5–0.7 percentage points below full precision, and a small recommender whose error matched full precision 3.
  • Auditor-in-a-Box. A reference implementation runs in Tinfoil confidential virtual machines 6. Its authors state that user data and plan execution in the demo are not actually secure, and that it has not been stress-tested by a counterparty 6.

Limitations

  • Verdict leakage. Abdelghafar and Kulp found that one-bit reports can reveal sensitive attributes 7. They propose designing the evidence itself to limit this 7.
  • Physical attacks on TEEs. Researchers interposing on the memory bus extracted a per-CPU Intel attestation key and forged TDX attestations 14, and a second team forged TDX attestation reports with an active interposer 13. Other attacks forged AMD SEV-SNP attestations, one through a DDR4 interposer and one from software alone on platforms without AMD's fix 15 16.
  • Stack trust. Compromise of the Docker daemon, host kernel or TEE stack breaks Cove's guarantees 2.
  • Scale. Frontier model inference typically needs the resources of several GPUs 8. Pour Demain's evaluation used one server with eight H200 GPUs, and the double-blind pilot one H100 9 10. ZkAudit was shown on image classifiers and a recommender model, not frontier-scale language models 3.
  • Process. Plan negotiation, false positives and appeals remain open problems 6.

For attesting that a declared safeguard ran on a single service, see Safeguard attestation.

Known flaws

Published flaws, with their severity, kind and status. How flaws are rated.

  • Released verdicts can leak information about private inputsSignificantDemonstrated attackOpen

    Even a one-bit result can reveal more than intended. Abdelghafar and Kulp used the published evaluation of Auditor-in-a-Box, whose output filter is meant to disclose at most one bit. Given only the valid or invalid decision on a new request, a simple predictor identified which of two request sets it came from (borderline or proxy requests versus ordinary research requests) with 72.4% balanced accuracy, against 50% without the decision. In a second experiment, which distinguished reinforcement-learning workloads from others, several candidate one-bit reports revealed exact-workload information.

    Sources: [7]
  • Memory-bus interposition extracts attestation keys and forges attestationsSignificantDemonstrated attackOpen

    With physical access to a server's DDR5 memory bus, researchers extracted a per-CPU Intel attestation provisioning key and forged TDX attestations. Against AMD SEV-SNP the same attack recovered a signing key used inside the virtual machine, not an AMD attestation key. By pairing forged TDX attestations with genuine H100 attestations relayed from rented hardware, they made a workload without TEE protection appear to run under GPU confidential computing. Intel, AMD and NVIDIA acknowledged the findings, and the researchers report that Intel and AMD consider interposer attacks out of scope. A second team forged attestation reports on an up-to-date Intel TDX platform with an active DDR5 interposer, DDRop. Other attacks have forged SEV-SNP attestation: Battering RAM with an interposer on DDR4 servers, and RMPocalypse from a malicious hypervisor on platforms without AMD's firmware fix. Cove's reference implementation trusts Intel TDX attestation roots, and the double-blind pilot ran on Intel TDX with an H100.

  • Guarantees depend on the host software stack and on reviewSignificantTheoretical argumentOpen

    Cove's developers state that compromise of the Docker daemon, host kernel or TEE stack breaks all guarantees. They also state that Docker policy alone cannot prove that guest code cannot generate a quote if the platform exposes quote instructions globally, and that compiled workflow bundles are hashed and reviewable but not signed by a publisher key.

    Sources: [2]
  • Completeness is not establishedSignificantTheoretical argumentOpen

    A confidential workflow proves facts about the records and models submitted to it. A governance analysis notes that an auditor also needs assurance that all activity is accounted for, since a host could start a second confidential virtual machine that uses a different model or bypasses monitoring.

    Sources: [8]
  • Zero-knowledge audits reveal model architectureMinorTheoretical argumentOpen

    ZkAudit keeps weights and data secret but reveals the model architecture, and it does not protect against data poisoning.

    Sources: [3]

Blockers

Technical detail

  • Cove object model. An artifact is a unit of private data with exactly one owner, encrypted locally before upload. A workflow is a directed acyclic graph with a canonical serialization, and its hash names the bundle. Compilation to per-node manifests is deterministic, and the manifest hash is the node's identity. Owners record allow rules that bind each artifact to approved manifest hashes 2.
  • Cove run and verification. At run time a node fetches and verifies upstream certificates, evaluates preconditions, obtains decryption keys only after attestation, runs the workload and emits a certificate whose TEE report data commits to the certificate-body hash. Report data is two SHA-256 hashes, of a fixed label and of a canonical payload. Verifiers start from a terminal certificate, the workflow bytes and the TEE vendor's roots, and check each certificate recursively 2.
  • ZkAudit. The provider publishes commitments to its dataset and weights, with a zero-knowledge proof that the committed weights result from training. It then answers each audit request by computing a function F privately and releasing F's output with a second proof 3. A copyright or demographic audit of a MobileNet v2 model on Flowers-102 cost $108 in total (about 10 cents per image), and a counterfactual audit of a recommender on MovieLens cost $8,456 3.
  • Minimal Information Disclosure. It chooses the evidence mechanism to minimize the conditional mutual information between protected properties and the released evidence, subject to a floor on information about the authorized target. One variant proves a linear projection with a Groth16 zk-SNARK 7.

Sources

  1. BS. Ding et al. (2026). Cove: Compositional Multi-Party Confidential Workflows for Verifiable AI Governance. ICML 2026 Workshop on Technical AI Governance Research. Source recordSupports: problem statement; framework; three applications; open-source implementation on Intel TDX via dstack · abstract (read via the ICML 2026 virtual poster page; the OpenReview PDF was not reachable)
  2. Bcovehub (2026). Cove: Compositional Multi-Party Confidential Workflows for Verifiable AI Governance (reference implementation). GitHub. Source recordSupports: object model; lifecycle; certificates; trust boundary; residual risks · README; docs/internal/architecture.md; docs/internal/security_model.md
  3. AS. Waiwitlikhit et al. (2024). Trustless Audits without Revealing Data or Models. 41st International Conference on Machine Learning (ICML 2024). Source recordSupports: ZkAudit protocol; models and datasets; accuracy; costs; assumptions; architecture disclosure; data poisoning · abstract; §5; Tables 1-4; limitations
  4. BC. Schnabl et al. (2025). Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments. ICML 2025 Workshop on Technical AI Governance. Source recordSupports: multi-party enclave audit protocol; transparency log; prototype and throughput; CPU versus GPU cost and slowdown; vendor trust · §3; §4; §5; Table 2
  5. BB. Penchas et al. (2026). Enabling Verifiably-Scoped Monitoring through Large Language Models and Trusted Compute. ICML 2026 Workshop on Technical AI Governance Research. Source recordSupports: plan-scoped monitoring protocol · abstract
  6. CR. Rinberg & B. Penchas (2026). Auditor-in-a-Box: Tools for Third-Party Auditing. LessWrong. Source recordSupports: plan definition; reference implementation; process problems; limitations · whole post
  7. BS. Abdelghafar & G. Kulp (2026). Privacy-Preserving AI Verification via Minimal Information Disclosure. arXiv. Source recordSupports: minimal information disclosure framework; one-bit leakage findings; Groth16 variant; limitations · abstract; introduction; Appendix A; Figure 5; limitations
  8. CGloria Z (2026). On TEEs for Privacy-Preserving Monitoring in AI Governance. MIRI Technical Governance Team. Source recordSupports: completeness and second-CVM gap; vendor root of trust; GPU TEE maturity; multi-GPU inference; treaty threat model · resource accounting; hardware auditability; physical attack surface
  9. BA. Trask et al. (2026). Double Blind Evals: Resolving the Dual Confidentiality Dilemma in AI Safety Auditing. Google DeepMind. Source recordSupports: double-blind evaluation pilot: participants, model, benchmark, GCP Confidential Space on H100 with Intel TDX, PySyft, mutual attestation checks, overhead figure cited from NVIDIA, uninspected model code, Google in the verification path, scaling to many-node clusters · abstract; architecture; limitations; future work
  10. CA. Tlaie Boria (2026). Confidential computing can enable better frontier AI auditing. Pour Demain. Source recordSupports: Pour Demain's interpretability evaluations of GLM-5.1 on Tinfoil (Intel TDX, eight H200 GPUs): enclave-bound tensors, bounded signed exports, overheads, single-session governance · whole post
  11. BGoogle Cloud (2026). Confidential Space overview. Google Cloud documentation. Source recordSupports: Confidential Space: multi-party roles; data released only to attested workloads; operator has no access; supported TEEs · overview
  12. BGoogle Cloud (2026). Confidential Space release notes. Google Cloud documentation. Source recordSupports: Confidential Space generally available, including on H100 GPUs from 2026-04-29 · release notes, 2023-03-28 and 2026-04-29
  13. AJ. De Meulemeester et al. (2026). DDRop: Active Memory Interposer Attacks on Confidential VMs by Dropping DDR5 Writes. 2026 ACM SIGSAC Conference on Computer and Communications Security (CCS '26). Source recordSupports: DDRop forges attestation reports on an up-to-date Intel TDX platform with an active DDR5 interposer · site summary
  14. AJ. Chuang et al. (2026). TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition. 2026 IEEE Symposium on Security and Privacy (SP). Source recordSupports: physical extraction of Intel attestation keys and SEV-SNP signing keys; forged attestations against NVIDIA GPU confidential computing; vendor acknowledgement and positions · project site summary; paper abstract and disclosure
  15. AJ. De Meulemeester et al. (2026). Battering RAM: Low-Cost Interposer Attacks on Confidential Computing via Dynamic Memory Aliasing. 47th IEEE Symposium on Security and Privacy (S&P 2026). Source recordSupports: Battering RAM forges SEV-SNP attestation with a DDR4 interposer
  16. AB. Schlüter & S. Shinde (2025). RMPocalypse: How a Catch-22 Breaks AMD SEV-SNP. 2025 ACM SIGSAC Conference on Computer and Communications Security (CCS '25). Source recordSupports: RMPocalypse forges SEV-SNP attestation from a malicious hypervisor
  17. BAMD (2025). SEV-SNP RMP Initialization Vulnerability (AMD-SB-3020). AMD product security bulletin. Source recordSupports: AMD firmware fixes for RMPocalypse (CVE-2025-0033)

Search

Full search page