Model identity attestation
Model identity attestation shows users, auditors and regulators that a provider is serving the model weights it committed to.
There are two routes. First, an attestation from a trusted execution environment shows that measured software enforced a hash commitment to the weights while the model ran. Second, a verifier holding the declared weights recomputes a sample of logged outputs, which can detect weights smuggled out in responses.
Tinfoil reports running the enclave route commercially with its open-source Modelwrap tool. For unpublished weights, the commitment alone shows only that the same weights are served each time. Research prototypes bind evaluations to the same hash, so results apply to the served model.
The enclave route inherits the weaknesses of TEEs, including published attacks that forge Intel TDX and AMD SEV-SNP attestations. Recomputation, tested on models of up to 30 billion parameters, needs trusted logging and randomness, and must tolerate numerical nondeterminism.
Tinfoil reports running the enclave route in a production service, but the independent attacks published so far left flaws open, including a critical flaw in the trusted execution environments that route relies on.
Rubric assessment
Assessed use: showing users that a service runs the declared model weights
- R1 met: designs and assumptions are published for both routes 1 2 4.
- R2 met through Tinfoil's Modelwrap chain. Its code is open source 16, it runs on NVIDIA H100, H200 or B200 GPUs with AMD SEV-SNP or Intel TDX (provider-reported) 13, and it has been reported on models of up to 554 GB 1. Rinberg et al. publish code and results on models from 3B to 30B parameters 4.
- R3 met on the provider's account: Tinfoil reports offering the feature in a production service 14 15.
- R4 not met, because the independent evaluations that exist left flaws open. TEE.fail used physical access to forge the Intel TDX attestation that the enclave route relies on. By pairing the forgeries with relayed H100 attestations, it made a vLLM proxy running outside TEE protection pass both the TDX and the NVIDIA confidential-computing checks 9. That flaw is critical and open. Kezins, at Delft University of Technology, attacked the exfiltration bound of Rinberg et al.'s recomputation check with chosen prompts. On a 30B mixture-of-experts model this raised leakage from 0.119 to 0.286 bits per token. The attack widens the covert channel and does not target the check that outputs match the declared model 5.
- Independent security evaluation of a deployed model-identity scheme that leaves no critical flaw open.
- Resistance of the enclave variant to physical attackers (see TEE remote attestation for AI workloads).
- Third-party verification for private models beyond consistency across requests.
- Tooling for audit-time checking of transparency records.
Assessed 2026-09-25 against rubric v1.1.
On this page
How it works
Two routes exist, and they can be combined.
Commit and attest
This route runs inside a TEE (TEE remote attestation for AI workloads). It closes a gap: attestation "measures launch state, not runtime state", while weights are read from disk after boot 1. Three designs bind the weights to the attestation:
- Tinfoil's Modelwrap packs the weights into a read-only image and computes a Merkle-tree root hash. It places that hash on the kernel command line, which the enclave measurement covers, and the kernel's dm-verity layer checks every block read against it 1. Clients compare the attestation's measurements with values that Tinfoil's build pipeline publishes to the Sigstore transparency log 14. For a public model, anyone can rebuild the image from the same Hugging Face commit and compare root hashes 1.
- PAL*M puts hashes of each operation's inputs, including the model, and of its outputs into the Intel TDX report, for single prompts and multi-turn sessions 2.
- Attestable Audits records the model hash when an audit runs. At inference time it checks that the served model matches, then returns each response with an attestation that links model, prompt, output and audit result 3.
Recompute and compare
This route uses sampled recomputation. A verifier that holds the declared weights scores logged outputs by the likelihood that each token was sampled from that model under a known seed. The estimators allow for legitimate numerical nondeterminism 4. Rinberg et al. frame the check as a way to catch weights hidden steganographically in responses, and it also shows whether outputs are consistent with the declared model 4. Deterministic inference is covered in Deterministic and bit-exact inference.
What it establishes
The two routes establish different things:
- Commit and attest can show that the bytes served match a commitment 1.
- Recompute and compare can show that logged outputs are consistent with the declared model's sampling procedure 4.
Neither shows what a private model can do, everything else that shapes behaviour, or properties that a weak hashing scheme leaves out:
- With unpublished weights, a user can confirm only that the same weights are served each time 1. An attested evaluation bound to the same hash can close part of that gap 3.
- Tinfoil measures the firmware, kernel, initrd and a configuration file as well as the weights 14. Gloria Z warns that unmeasured runtime flags can undermine integrity 6. In a pre-launch audit of WhatsApp's TEE-based Private Processing, Trail of Bits found such a gap: configuration files with environment variables were loaded after the measurement, so an insider could have loaded malicious code without invalidating the attestation. Meta fixed it before launch 7 8.
- Gloria Z also warns that a scheme that sequences values loosely, or checks the weights too early in loading, could let two models with significantly different properties share a hash "without breaking the hash function itself" 6.
Threat model
The enclave route trusts the hardware vendor and treats the host software as hostile. Tinfoil states "You trust the hardware manufacturer", but not the hypervisor or host OS 13. PAL*M excludes side channels and physical attacks 2.
The recomputation route assumes the following 4:
- the verification code is trusted;
- every input, output and seed is logged correctly;
- the attacker can neither predict nor manipulate which messages are sampled for verification;
- the verifier sees prompts and responses and holds the weights.
Evidence
- Tinfoil. Modelwrap is published under an MIT license 16. Tinfoil reports a storage overhead of 0.8%, build times of up to 13 min 25 s for a 554 GB model, and slower cold loading but no inference slowdown 1.
- PAL*M. It reports inference attestation on an H100 across three models of 3.8 to 8 billion parameters. Its added time was 3.8–11.4% of total run time in multi-turn sessions and 45.5–66.4% for single prompts. The authors plan to release the code after peer review 2.
- Attestable Audits. Its protocol binds inference to the audited model hash. The reported evaluation covers the audit step, for a 4-bit Llama-3.1-8B on CPU-only AWS Nitro Enclaves 3.
- Rinberg et al. They tested Llama-3.1-8B, Llama-3.2-3B and two Qwen mixture-of-experts models, and publish their code. On the 30B model under benign prompt traffic, the detector cut exfiltratable information to under 0.5% at a false-positive rate under 0.01% 4. An independent study found that prompt control roughly doubles leakage per token 5.
- Apple Private Cloud Compute. Apple reports that devices send requests only to servers that attest to software in a public transparency log, and that model assets share the code's integrity protection 17.
- Verde. Gensyn reports that its refereed-delegation system shows a delegated output came from the declared model and data, by re-running disputed operations with bitwise-reproducible kernels 18. The guarantee holds if at least one of the compute providers is honest 19.
Limitations
- Inherited TEE attacks. Using physical access, TEE.fail forged Intel TDX attestations and, by pairing them with relayed H100 attestations, passed a workload outside TEE protection 9. Battering RAM and RMPocalypse forged AMD SEV-SNP attestations, the second without physical access 10 11. Tinfoil's documentation acknowledges physical-access and side-channel risks 13.
- Audit-time checking. Tinfoil does not provide a supported tool for querying its transparency records at audit time 15.
- Recomputation slack. Nondeterminism concentrates at a few token positions, and very slow leaks remain possible 4. An adversary who controls prompts widens the leak 5.
- Recomputation dependencies. The check depends on complete and trusted logging, and on the verifier holding the weights 4.
Known flaws
Published flaws, with their severity, kind and status. How flaws are rated.
Underlying attestation can be forged or relayed
The enclave variant is only as sound as the attestation. With physical access and root privileges, TEE.fail extracted a per-CPU Intel attestation signing key and forged TDX attestations. Pairing the forgeries with genuine H100 attestations relayed from a rented server, the researchers made a workload outside TEE protection pass both checks 9. The authors report that Intel and AMD treat interposer attacks as outside their threat models. They advise users to run confidential workloads only on physically secure servers 9. Other independent research forged AMD SEV-SNP attestations. Battering RAM did so with a DDR4 memory interposer, and RMPocalypse did so from a malicious hypervisor without physical access 10 11. AMD reports firmware fixes for RMPocalypse 12. Tinfoil's documentation lists physical attacks among the limitations of enclaves 13.
Launch-state attestation does not by itself cover weights loaded later
Attestation measures launch state, and weights are read from disk after boot. A signature checked at load time does not stop a malicious hypervisor from altering the disk afterwards 1. Tinfoil reports mitigating this with dm-verity checks on every read 1. Unmeasured runtime configuration remains a general risk 6.
Recomputation depends on trusted logging and randomness, and its tolerance leaves a covert channel
The recomputation variant assumes that every input, output and seed is logged correctly, and that the attacker can neither predict nor manipulate which messages are sampled for verification. Legitimate nondeterminism concentrates at a few token positions, and slow leaks within the tolerated slack remain possible 4. An independent study showed that an adversary who controls the prompts roughly doubles the bits leaked per token, reducing the exfiltration slowdown from 146–254 times under benign prompts to 60–118 times. The attack targets the exfiltration bound, not the check that outputs match the declared model 5.
Blockers
Attestation that resists physical attackers, for the enclave variant.
Numerical nondeterminism limits how tightly recomputation can pin down the model and sampling.
The recomputation variant needs the verifier to hold the declared weights.
Technical detail
- Modelwrap build (Tinfoil's description). The tool downloads a pinned model revision, normalizes its directory structure and builds an EROFS image. It computes a dm-verity root hash with veritysetup. The Merkle tree gives a 32-byte commitment to, for example, 140 GB of weights 1.
- Binding and enforcement. The root hash is passed on the kernel command line, which the enclave measurement covers. dm-verity then checks every block read against it and returns an error on any mismatch 1.
- Reported costs. Tinfoil reports a hash-tree overhead of about 0.8% of image size, and build times from 5 s for a 549 MB model to 13 min 25 s for a 554 GB model. Cold-cache loading takes about 80% longer, and inference is unaffected once the weights are in GPU memory 1.
- PAL*M inference attestation. PAL*M sets the Intel TDX REPORTDATA field to the concatenation of the operation, a challenge, and hashes of the inputs and outputs. It supports single-prompt and multi-turn inference attestations. Across three models of 3.8 to 8 billion parameters on an H100, PAL*M's added time was 3.8–11.4% of total run time for multi-turn sessions and 45.5–66.4% for single prompts, so a single attested prompt took 1.8 to 3 times as long as without PAL*M 2.
- Fixed-seed sampling likelihood. Rinberg et al. score the likelihood that each token was honestly sampled from the claimed model under a known seed, with estimators for inverse-probability-transform and Gumbel-max samplers. On a 30B mixture-of-experts model under benign prompt traffic, the check limits exfiltratable information to under 0.5% at a false-positive rate under 0.01%, a slowdown of more than 200 times for the adversary 4. An independent study found that an adversary who controls the prompts roughly doubles leakage per token, reducing the slowdown to 60–118 times 5.
Sources
- CTinfoil Team (2026). How Tinfoil Proves Exactly What Model Is Running. Tinfoil. Source recordSupports: Modelwrap design, launch-state problem, signing comparison, private models, overheads (provider-reported) · sections on the challenge, the three phases, performance, private models
- BP. Chantasantitam et al. (2026). PAL*M: Property Attestation for Large Generative Models. arXiv. Source recordSupports: inference attestation binding hashes to TDX report; overheads · §4, Table 6
- BC. Schnabl et al. (2025). Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments. ICML 2025 Workshop on Technical AI Governance. Source recordSupports: audit-to-inference model-hash binding; prototype evaluation of the audit step · §3, Algorithms 1-3, §5
- BR. Rinberg et al. (2025). Verifying LLM Inference to Detect Model Weight Exfiltration. arXiv. Source recordSupports: recomputation-based verification, assumptions, results, limitations, code · Abstract; §1, §3.3-3.4, §4.2, §5-§7
- BN. Kezins (2026). Adversarial Entropy Inflation Against Gumbel-Based Inference Verification. arXiv. Source recordSupports: independent prompt-control attack that widens the exfiltration bound; 0.119 to 0.286 bits per token on the 30B MoE model · Abstract; Table 1
- CGloria Z (2026). On TEEs for Privacy-Preserving Monitoring in AI Governance. MIRI Technical Governance Team. Source recordSupports: measurement incompleteness; hashing-scheme warning
- CTrail of Bits (2026). What we learned about TEE security from auditing WhatsApp's Private Inference. Trail of Bits blog. Source recordSupports: Trail of Bits audit of WhatsApp Private Processing: environment variables loaded after the measurement (TOB-WAPI-13) and Meta's fix
- BTrail of Bits (2025). Meta WhatsApp Private Processing (security review). Trail of Bits publications library. Source recordSupports: the review's finding that CVMs could be compromised through environment-variable injection
- AJ. Chuang et al. (2026). TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition. 2026 IEEE Symposium on Security and Privacy (SP). Source recordSupports: Intel TDX attestation forgery and H100 attestation relay to a vLLM proxy outside TEE protection · Abstract; §1.1, §8.3
- AJ. De Meulemeester et al. (2026). Battering RAM: Low-Cost Interposer Attacks on Confidential Computing via Dynamic Memory Aliasing. 47th IEEE Symposium on Security and Privacy (S&P 2026). Source recordSupports: SEV-SNP attestation breach with a DDR4 interposer (Battering RAM) · Abstract; site FAQ
- AB. Schlüter & S. Shinde (2025). RMPocalypse: How a Catch-22 Breaks AMD SEV-SNP. 2025 ACM SIGSAC Conference on Computer and Communications Security (CCS '25). Source recordSupports: software-only SEV-SNP attestation forgery by a malicious hypervisor (RMPocalypse) · Abstract; site
- BAMD (2025). SEV-SNP RMP Initialization Vulnerability (AMD-SB-3020). AMD product security bulletin. Source recordSupports: AMD firmware fixes for RMPocalypse (vendor-reported) · Mitigation tables
- BTinfoil (2026). A primer on secure enclaves. Tinfoil documentation. Source recordSupports: Tinfoil hardware, trust model and documented limitations (provider-reported)
- BTinfoil (2026). Backend infrastructure. Tinfoil documentation. Source recordSupports: measured boot chain, Sigstore publication, production model volumes (provider-reported)
- BTinfoil (2026). How verification works in Tinfoil. Tinfoil documentation. Source recordSupports: production deployment; no supported audit-time tool (provider-reported) · In-band vs. out-of-band verification
- BTinfoil (2026). modelwrap: Reproducible dm-verity read-only image of Huggingface models. GitHub. Source recordSupports: open-source implementation, release v0.3.0
- CApple Security Engineering and Architecture (SEAR) (2024). Private Cloud Compute: A new frontier for AI privacy in the cloud. Apple Security Research blog. Source recordSupports: Apple PCC: integrity protection of code and model assets; attestation against a public transparency log (provider-reported)
- CO. Ersoy (2025). Verde Verification System In Production. Gensyn research blog. Source recordSupports: Gensyn's statement that Verde shows outputs came from the declared model and data (provider-reported)
- BA. Arun et al. (2025). Verde: Verification via Refereed Delegation for Machine Learning Programs. arXiv. Source recordSupports: Verde's guarantee holds if at least one compute provider is honest · Abstract