The declared model is the one being served
The claim is that outputs delivered to users or auditors come from the specific model, weights and configuration the provider declared.
Evaluations, audits and agreements often apply to one specific model. If a provider could evaluate one model and serve another, such as a cheaper, quantized or modified version, those checks would say little about what users receive. It is a positive claim that can be tested directly, but three problems make it hard.
Numerical nondeterminism means honest recomputation does not match exactly. The verifier usually cannot see the weights, which are commercially or strategically sensitive. And the evidence must come from the actual serving system rather than a separate test instance.
Approaches include statistical or exact recomputation of sampled outputs, hardware attestation of the loaded weights, and zero-knowledge proofs. They trade off cost, trust in hardware vendors and confidentiality.
The served model can be checked directly on its outputs, and several mechanisms for it are demonstrated (R2) or in production (R3). None is deployment-ready (R4): independent security evaluations are either missing or, for trusted execution environments, found a critical flaw.
For example, users of a hosted model may want to know that the model answering them is the one an auditor evaluated.
- Model identity attestation (R3) extends TEE remote attestation (R3) from the software an enclave booted to the weights it loads. Confidential multi-party verification (R2) runs the audit in an enclave and binds its result to the model's hash 10.
- Attestable Audits (R2) joins the two steps. At inference, the enclave checks the served model's hash against the audited one and returns each response with an attestation that links model, prompt, output and audit result 10. Its reported evaluation covers only the audit step, on CPU-only enclaves with a 4-bit, 8-billion-parameter model 10. Tinfoil's Modelwrap (R3) covers serving alone, in what Tinfoil reports is a production service 12.
- The evidence is narrower than what users want. Attested enclaves report that the weights that answered have the same hash as the weights that were audited. For a public model, Modelwrap lets anyone rebuild the hash from the published weights. For a private model, users see only the hash 11.
- The chain trusts the TEE vendors' attestation keys. With physical access, researchers forged Intel TDX attestations and paired them with relayed H100 attestations, so that a system outside TEE protection passed both checks 13. Sampled recomputation (R3), as in DiFR, avoids hardware trust but needs a trusted copy of the weights and evidence tied to the production serving path 6. Zero-knowledge proofs of inference (R2) keep the weights private without trusting hardware, but proving remains expensive 6.
On this page
Mechanisms
- Enables exact-match recomputation checks that the declared model, weights and software setup produced the outputs.
- Core purpose: responses come from the declared weights.
- Checks that sampled recorded outputs are consistent with the declared model, precision and sampling settings.
- Attests the software stack that produced responses, and the model too when paired with a weight commitment (see Model identity attestation).
- Clients check that the served weights match a committed root hash.
- R3TOPLOCprimaryChecks that the provider produced outputs with the claimed model weights, prompt and precision.
- R3Verde and RepOps (Gensyn)primaryA client learns that a delegated inference output came from the declared model and input, if at least one provider is honest 1 2.
- R2Attestable AuditsprimaryUsers can check that the model answering them is the audited one.
- Binds audit or capability-evaluation results to the model that is served, without revealing weights 1 4.
- Checks that outputs are consistent with the declared model, precision and sampling configuration.
- R2EZKLprimaryProves an output follows from a committed model. South et al.'s results reach about a million parameters 2.
- Both parties approve the attested workload that runs a private evaluation against the model and prompts they submit 1.
- Binds each proven output to committed weights and a public architecture.
- R2zkLLMprimaryProves an output follows from committed weights and a public architecture.
- Attestable reports proving y = F(W, x, r) for committed weights W.
- R3Apple Private Cloud ComputesupportingAttests the software release that served a request. Apple reports that model assets share the code's integrity protection 1.
- Makes exact-match recomputation of served outputs possible when the verifier runs the same model, engine and hardware 6.
- R2Safeguard attestationsupportingProperty and audit attestations bind responses to a measured model 4 5.
- Recomputation checks sampled packets against the declared model.
- Deployment only to approved flexHEG devices, and verification of evaluation scores 1.
- Screening checks that the model is on an agreed whitelist.
- R1Network taps and certifierssupportingReplaying challenged records with the declared model checks which model produced outputs 1 4.
- The compute sanctum checks resident weights against reference measurements before serving.
- R1SASH confidential network loggersupportingRecomputation uses another copy of the declared model 1 2.
Why it matters
Checks on a model's capabilities or safety apply to the model that was checked. Several parties depend on that link:
- RAND's framework asks verifiers to confirm that declared inference is declared accurately, so that the prover actually did the claimed deployment 1. It then asks them to confirm that deployed models have the required properties, for example through evaluations at regular intervals 1. Both steps assume the evaluated model is the served model. A low-trust system overview identifies running approved models for prohibited uses as the most difficult violation to defend against, and aims to deploy only approved models 6. The Oxford Martin report includes appendices on model fingerprint attestation and on "device-model mating" with an encrypted model 5.
- The DiFR authors note that providers and their customers increasingly need to verify that inference is performed correctly, without errors or tampering 2. Their Token-DiFR method detected 4-bit quantization of a model with AUC above 0.999 within 300 output tokens 2. In an audit of commercial inference APIs for four Llama models in summer 2024, Gao, Liang and Guestrin found that 11 of 31 endpoints served a different output distribution from the published reference weights 16.
Why it is hard
- Statistical tests on an API's outputs need no cooperation from the provider, but they have limits 16. Gao, Liang and Guestrin's test reached a median statistical power of 77.4% against a range of distortions, using an average of 10 samples per prompt 16. Cai and colleagues find that tests on text outputs are query-intensive and fail against subtle substitutions, and that tests on log probabilities are defeated by inference nondeterminism in production 17.
- Re-running the same inference often gives slightly different results because of benign numerical variation, which makes it hard to tell legitimate variation from substitution 2. In LLM serving a major cause is that kernels are not invariant to batch size, which varies with server load 3. Statistical tests tolerate this noise 2. Alternatively, Cankaya reports that inference can be reproduced bit-exactly across several NVIDIA GPU variants in software, given enough recorded information about the original run 4.
- Recomputation requires the weights, which a provider or state may not reveal 6. One system design keeps weights cryptographically committed and uses them only inside air-gapped auditing environments 6. Zero-knowledge proofs avoid revealing weights at all. zkLLM reports proving one 2,048-token forward pass of a 13-billion-parameter model in under 15 minutes, with proofs under 200 kB 7, but the low-trust overview describes zero-knowledge proof computation as still expensive 6. Trusted execution environments can attest which software and data were loaded 9. PAL*M reports property attestation on confidential computing hardware (Intel TDX with NVIDIA H100) at under 11% overhead for common operations 8.
- Evidence about a test instance says nothing about production unless it is tied to the serving hardware and time. The low-trust design aims to identify each forward pass uniquely and attribute it to the hardware and time it was processed 6. TEE-based binding relies on the vendor's keys: whoever holds the hardware's attestation key can produce valid reports, and every component that influences inference must be covered by launch measurement 9. A published memory-bus interposition attack, which needs physical access, extracted a per-CPU Intel attestation key and forged Intel TDX attestations 13. Other attacks forged AMD SEV-SNP attestations, one through a DDR4 interposer and one from software alone on platforms without AMD's fix 14 15.
Sources
- BM. Baker et al. (2025). Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment. RAND Corporation. Source recordSupports: Subgoal 1.A (declared uses declared accurately, including inference) and 1.B (required properties; deployed models evaluated at intervals) · §3.2
- BA. Karvonen et al. (2025). DiFR: Inference Verification Despite Nondeterminism. ICML 2026 Workshop on Technical AI Governance Research. Source recordSupports: need to verify inference; nondeterminism; Token-DiFR detects 4-bit quantization with AUC > 0.999 within 300 tokens · abstract
- CH. He & Thinking Machines Lab (2025). Defeating Nondeterminism in LLM Inference. Thinking Machines Lab: Connectionism. Source recordSupports: batch-size dependence as a cause of inference nondeterminism · batch invariance section
- BN. Cankaya (2026). Bit-Exact AI Inference Verification Without Performance Tradeoffs. ICML 2026 Workshop on Technical AI Governance Research. Source recordSupports: bit-exact reproduction across GPU variants given recomputation data · abstract
- BB. Harack et al. (2025). Verification for International AI Governance. Oxford Martin AI Governance Initiative. Source recordSupports: model fingerprint attestation; device-model mating with an encrypted model · Appendix K (p. 157); Appendix L.4 (p. 159)
- BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: whitelisted models for blacklisted uses; attributing forward passes to hardware and time; committed weights in auditing environments; ZKP cost · verification goals; architecture; open problems
- AH. Sun et al. (2024). zkLLM: Zero Knowledge Proofs for Large Language Models. 2024 ACM SIGSAC Conference on Computer and Communications Security (CCS 2024). Source recordSupports: zkLLM proves one 2,048-token forward pass of a 13B-parameter model in under 15 minutes with proofs under 200 kB, keeping parameters private · abstract; §8 Table 1
- BP. Chantasantitam et al. (2026). PAL*M: Property Attestation for Large Generative Models. arXiv. Source recordSupports: property attestation on Intel TDX + NVIDIA H100 with under 11% overhead for common operations · abstract
- CGloria Z (2026). On TEEs for Privacy-Preserving Monitoring in AI Governance. MIRI Technical Governance Team. Source recordSupports: attestation-key holder can produce valid reports; side-channel and physical attacks; measurement coverage · Limitations
- BC. Schnabl et al. (2025). Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments. ICML 2025 Workshop on Technical AI Governance. Source recordSupports: attestation linking model, audit result, prompt and response in a TEE; reported evaluation covers the audit step on CPU-only enclaves with a 4-bit 8B model · abstract; inference protocol; §5
- CTinfoil Team (2026). How Tinfoil Proves Exactly What Model Is Running. Tinfoil. Source recordSupports: public models' hashes can be rebuilt; private models expose only the hash (provider-reported)
- BTinfoil (2026). How verification works in Tinfoil. Tinfoil documentation. Source recordSupports: Modelwrap chain in Tinfoil's production service (provider-reported) · In-band vs. out-of-band verification
- AJ. Chuang et al. (2026). TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition. 2026 IEEE Symposium on Security and Privacy (SP). Source recordSupports: physical memory-bus interposition extracts a per-CPU Intel attestation key and forges TDX attestations · abstract; §1.1
- AJ. De Meulemeester et al. (2026). Battering RAM: Low-Cost Interposer Attacks on Confidential Computing via Dynamic Memory Aliasing. 47th IEEE Symposium on Security and Privacy (S&P 2026). Source recordSupports: Battering RAM forges SEV-SNP attestation with a DDR4 interposer
- AB. Schlüter & S. Shinde (2025). RMPocalypse: How a Catch-22 Breaks AMD SEV-SNP. 2025 ACM SIGSAC Conference on Computer and Communications Security (CCS '25). Source recordSupports: RMPocalypse forges SEV-SNP attestation from a malicious hypervisor
- AI. Gao et al. (2025). Model Equality Testing: Which Model Is This API Serving?. International Conference on Learning Representations (ICLR 2025). Source recordSupports: model equality testing: median 77.4% power with about 10 samples per prompt; 11 of 31 Llama API endpoints in summer 2024 served distributions different from the reference weights · abstract
- BW. Cai et al. (2025). Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs. arXiv. Source recordSupports: output tests query-intensive and fail against subtle substitutions; log-probability tests defeated by inference nondeterminism · abstract