Claims

A claim is something one party wants to verify about another party's AI hardware or software, such as "this cluster only runs inference". Each entry gives the claim, where its verification stands, and the mechanisms that address it.

Compute stock is at most a declared amount

Negative claim · 1 primary, 2 supporting

A party holds no more AI-relevant compute, counted in chips or equivalent capacity, than the total it has declared.

State of verification

No mechanism can yet bound a party's chip stock, and every approach is proposed (R1). Chip tracking would reach new production far better than chips already in circulation, partly because the chip supply chain is concentrated 1 2.

Chip registries and manufacturing records (R1) would follow each chip from the fab to its owner, so that inspectors can check a sample against the declared records 1. Remote detection of data centres (R1) estimates the power capacity of large facilities from equipment visible outside, and public estimates exist for known large facilities 7. Performance throttling and licensing (R1) would cap the work that declared chips can do 8. Beyond these estimates, only designs and policy analyses are public, and no chip registry has been built for verification.

Millions of AI-relevant chips already exist with no central tracking 4, and domestic chip manufacture and older chips are listed as evasion routes 5. Draft agreements therefore pair technical measures with intelligence, inspections and whistleblowers 3.

Full record →

Chips are where they are declared to be

Positive claim · 2 primary, 3 supporting

Specific AI chips are physically located at the sites a party has declared, throughout the declared period.

State of verification

Chip location concerns known devices and can be checked positively, so it is one of the more tractable claims, but every publicly described scheme is still proposed (R1).

Chip location verification (R1) times a chip's signed replies to trusted servers, so that signal delay bounds its distance from them 1 2. Lucid's sovereignty certificates (R1) are a draft specification of this approach 10, and chip registries (R1) supply the declared locations to test. Guarantee processors (R1) could automate checks of approximate chip location, and their designers want them to be retrofittable to existing chip and server designs 11.

An IAPS issue brief reports a single result from a rudimentary prototype on NVIDIA H100 chips, a landmark in Singapore bounding a chip in Singapore to within 300 miles 7, and no systematic measurements, error rates or code have been published. NVIDIA has said that it is developing delay-based location verification using its own servers 8, but it has published no design or results, and its announced fleet-management software is opt-in 9.

The chip's private key must not be extractable, or another device can answer for it 2. Wasil and colleagues see location tracking as limited to newly produced chips 3, while Brass and Aarne expect that the H100's trusted execution environment could implement it 6.

Full record →

Declared hardware is idle or shut down

Negative claim · 2 primary, 3 supporting

Specified AI chips or facilities are not performing computation, or are powered off, throughout a declared period.

State of verification

Idleness is one of the more approachable negative claims, because computing needs power and leaves physical traces. On-chip telemetry and timed challenges are demonstrated (R2) as indicators of GPU activity, not as proofs that all declared hardware is idle.

Power draw, knowledge of on-site generation and possibly thermal imaging could show that a facility's chips are unpowered 1, but energy monitoring is unproven in practice and open to masking 2. On-chip telemetry (R2) and timed challenges (R2) could show whether a declared chip is busy, and GPU timing and memory measurements correlate with compute activity even when host and device are untrusted 7. Licensing and throttling (R1) would make chips refuse or slow work once a licensed budget is spent 5. Proofs of useful work (R1) instead keep declared hardware provably busy with agreed work 8.

A dark facility shows only that the hardware inside it is idle, so the claim depends on knowing where the declared chips are (Chips are where they are declared to be). Where chips must stay powered for permitted work, the claim becomes a bound on use (This compute runs inference, not training, A training run stayed within declared limits).

Full record →

This compute runs inference, not training

Mixed claim · 9 primary, 12 supporting

A declared cluster is used only to run existing models to produce outputs, and not to train new or more capable models.

State of verification

Declared inference can be checked with components that are demonstrated (R2) or in production (R3), but the absence of training cannot yet be verified. That half rests on full-stack designs (AI 2040 inference-only verification stack, Low-trust AI compute verification system overview, SASH confidential network logger) that are proposed (R1), and a team building the components reports nothing past a proof-of-principle prototype 13.

For example, a verifier may want to know that a cluster declared for inference is not training a new model.

  • Network taps (R1) record the cluster's front-end traffic with its users, and sampled recomputation (R3) re-runs sampled requests on the declared model. DiFR tolerates numerical noise in these re-runs on open-weight models of 8 to 30 billion parameters 11.
  • SASH's network logger (R1) is a public prototype of both. It passes every request through a logger and re-runs it on a separate cluster, with a 270-million-parameter model and no stated adversary 17.
  • This covers the positive half at most. It shows that sampled outputs match the declared model. A tap on the cluster's external links does not stop covert workloads. It aims only to stop their results leaving over those links 18.
  • The negative half needs the rest of the cluster accounted for. Training traffic runs on back-end fabric that is harder to tap 5. Bandwidth limits (R2 for software monitoring) between groups of chips target that fabric, since distributed training must exchange gradients 6. Attestable proposes proofs of useful work (R1) to keep declared hardware busy with approved or protocol-defined work, leaving little spare capacity 15.
  • Known routes remain. Reinforcement-learning rollouts are inference, so declared servers could generate them while hidden compute updates the model 16. In one scenario, JoshC estimates that more than 95% of computation must be accounted for to constrain that strategy 16. Workload classification (R2) is a lighter alternative. It detects training with 98.2% accuracy on its own corpus, but 43–87% on the most challenging disguised workloads held out from its training, and current GPUs lack the protections its telemetry needs to be trusted 8 14.

Full record →

The declared model is the one being served

Positive claim · 15 primary, 9 supporting

Outputs delivered to users or auditors come from the specific model, weights and configuration the provider declared, not from a substitute.

State of verification

The served model can be checked directly on its outputs, and several mechanisms for it are demonstrated (R2) or in production (R3). None is deployment-ready (R4): independent security evaluations are either missing or, for trusted execution environments, found a critical flaw.

For example, users of a hosted model may want to know that the model answering them is the one an auditor evaluated.

  • Model identity attestation (R3) extends TEE remote attestation (R3) from the software an enclave booted to the weights it loads. Confidential multi-party verification (R2) runs the audit in an enclave and binds its result to the model's hash 10.
  • Attestable Audits (R2) joins the two steps. At inference, the enclave checks the served model's hash against the audited one and returns each response with an attestation that links model, prompt, output and audit result 10. Its reported evaluation covers only the audit step, on CPU-only enclaves with a 4-bit, 8-billion-parameter model 10. Tinfoil's Modelwrap (R3) covers serving alone, in what Tinfoil reports is a production service 12.
  • The evidence is narrower than what users want. Attested enclaves report that the weights that answered have the same hash as the weights that were audited. For a public model, Modelwrap lets anyone rebuild the hash from the published weights. For a private model, users see only the hash 11.
  • The chain trusts the TEE vendors' attestation keys. With physical access, researchers forged Intel TDX attestations and paired them with relayed H100 attestations, so that a system outside TEE protection passed both checks 13. Sampled recomputation (R3), as in DiFR, avoids hardware trust but needs a trusted copy of the weights and evidence tied to the production serving path 6. Zero-knowledge proofs of inference (R2) keep the weights private without trusting hardware, but proving remains expensive 6.

Full record →

Declared safeguards were applied during inference

Positive claim · 1 primary, 8 supporting

Specified safety measures, such as input filters, output checks or monitoring, actually ran on the requests a deployed model served.

State of verification

Safeguard application has been shown only in small research prototypes, which can show that a declared safeguard ran, not that it works.

Safeguard attestation (R2) runs the safeguard inside a trusted execution environment, whose hardware signs a measurement of its code with a commitment to each input and response 8. It builds on TEE remote attestation (R3) and on knowing which model is served (The declared model is the one being served). The proof-of-guardrail prototype demonstrates it on CPU enclaves in AWS and calls its guardrail model through an external API 8. In the authors' tests it detected modified guardrail code, attestations and responses 8. Confidential multi-party verification (R2) can limit monitoring to a plan both parties sign 9, and Auditor-in-a-Box demonstrates such plan-scoped monitoring, though its authors state that the demo's user data and plan execution are not actually secure 10.

Every component that influences inference must be covered by the launch measurement 4, and the prototype attests only the responses for which it offers attestation, so coverage of all traffic is not shown 8. Where hardware trust is unavailable, the sources fall back on inspections, audits and personnel-based layers 1 2 3.

Full record →

A training run stayed within declared limits

Mixed claim · 3 primary, 8 supporting

A declared training run used no more compute than permitted and had its declared properties, such as data, hyperparameters and resulting weights.

State of verification

Proof-based checks of training remain far below frontier-scale runs. Proof-of-learning and zero-knowledge proofs of training are both demonstrated (R2); hardware enforcement remains proposed (R1).

Proof-of-learning and transcript verification (R2) re-runs sampled training segments between logged weight snapshots 1, and zero-knowledge proofs of training (R2) prove that training followed a committed specification without revealing data or weights 13. Guarantee processors such as flexHEG and licensing (both R1) could enforce compute limits in hardware 10 11.

Choi and colleagues report proof-of-training-data experiments on language models of up to 1 billion parameters 12, and Kaizen proves training iterations of a 10-million-parameter image model at about 15 minutes of proving per iteration 13. VeriLoRA proves individual low-rank fine-tuning steps on language models of up to 13 billion parameters 14.

Published attacks spoof the original proof-of-learning protocol, and the attack's authors argue that a provably robust version needs a better understanding of deep-learning optimisation 8. Governance analyses have judged zero-knowledge proofs impractical at frontier scale, and one 2026 proposal argues this is a limit of current approaches, not a fundamental one 7. A compute limit bounds a run only if all the chips used are known, so this claim depends on Compute stock is at most a declared amount and There is no undeclared relevant compute.

Full record →

Communication between compute groups is bounded

Negative claim · 2 primary, 4 supporting

Data flowing between specified groups of chips, or out of a facility, stays below a declared rate, so the groups cannot jointly run large workloads.

State of verification

A software bandwidth monitor has been demonstrated (R2). It runs on the operator's nodes, and its authors say the measurements are trivially spoofable 13. No bandwidth cap that a verifier can check has been demonstrated publicly.

Bandwidth limits and compartmentalization (R2 for software monitoring) proposes to cap or remove links between accelerator groups. Its pod-cap design aims to allow inference tokens while restricting training gradients, provided the served model and routing fit within each pod 1 15. It needs no access to a facility's code 1. Hardware versions include RAND's fixed-set design 4 and guarantee processors (R1), and the AI 2040 stack (R1) removes back-end networking 10. Side-channel suppression (R1) addresses paths outside the network 7. Tamper evidence (R2) would protect the enforcing devices, and bounding unexplained information (R2) limits what the permitted bandwidth can carry.

Lucid Computing's pod-cap design is "not yet implemented or red-teamed" 11, and Amodo's rate limiting on 400G links was for weight security with a cooperating operator 12.

Copper scale-up links and encrypted interconnects are hard to observe 6, and the designs aim to cut side channels to low rates, not to eliminate them 5 7. Sastry and colleagues note that decentralised training, if it becomes more viable, might undermine the detectability of training runs 2, and training methods that need less communication could likewise erode a bound that separates training from inference today.

Full record →

Model weights have not left the facility

Negative claim · 2 primary, 7 supporting

No copy of specified model weights has left a designated facility through networks, physical media or other channels.

State of verification

Only the network channel has a demonstrated (R2) bound, and only for inference outputs. Physical side channels rely on proposed (R1) designs, and insiders and physical media are left to facility security 1 and personnel measures such as whistleblowing 7.

For example, a verifier may want to know that a facility serving a model through an API has not let the weights out inside its responses.

  • Network taps (R1) record what the facility sends, and sampled recomputation (R3) re-runs sampled requests on a trusted copy of the model with the same random seed 2. Bounding unexplained information (R2) then scores how plausible each returned token is under honest sampling, which bounds how fast hidden bits can pass unnoticed 2.
  • Rinberg and colleagues published code for this check. On a 30-billion-parameter model under benign prompts, it cut the information an attacker could hide to under 0.5%, at a false-positive rate below 0.01% 2.
  • This establishes a rate limit on one channel, not that no copy left. In an independent study, an attacker who controlled the prompts roughly doubled the leak per token 3. One estimate still leaves about 40 Mbit/s of covert output for a 200,000-GPU inference cluster at full load, assuming about 0.1 bits per token survive replay checks 12. Deterministic inference (R3) would remove the tolerance that hidden data exploits 10.
  • The bound holds only if checked outputs are the only way out. Caps on external bandwidth (Communication between compute groups is bounded) limit the total that can leave 9. Side-channel suppression (R1) is designed to hold physical leakage below kilobits per second, not to zero 5. RAND's secure inference data centre (R1) is designed to protect weights against state-backed attackers but does not describe how an outside party would verify it 11.

Full record →

There is no undeclared relevant compute

Negative claim · 1 primary, 7 supporting

A party controls no AI-relevant computing hardware or facilities, above an agreed threshold, beyond those it has declared.

State of verification

No single verification method is foolproof for this broadest of negative claims 5, so draft agreements pair technical measures with intelligence, challenge inspections and whistleblowers 6. The only mechanism mapped primarily to this claim, remote detection of data centres, is proposed (R1), as are the chip registries and location checks that support it. RAND splits the claim into undeclared use of declared clusters, which reduces to claims such as Declared hardware is idle or shut down and This compute runs inference, not training, and undeclared clusters 1.

One strategy, which Scher and Thiergart favour, makes the declared stock complete from the start through chip registries and location verification 4. The other searches for what was missed, through remote detection of data centres and other national technical means 5. Satellite imagery, permits and utility filings already track the construction of known large facilities, but automated data-centre detection remains primarily conceptual 9. Proofs of useful work would leave declared hardware little spare capacity, but cannot find a facility that was never declared 10.

Chips sold before tracking began may not be locatable 3, facilities can be hidden underground or camouflaged 5, and it is unclear how small undeclared compute can be and still matter 1 7.

Full record →

Search

Full search page