Whole-workload recomputation (reproducible packets)

R1Proposed

The AI 2040 verification plan proposes that every AI workload in a monitored facility be organized into discrete, reproducible "packets" that a verifier's recomputation server can see.

The verifier re-runs a random sample of packets to check that they were computed as declared. The plan states that smaller packets raise the chance of catching a rogue workload. As of September 2026, no implementation of whole-workload packets has been published.

The nearest prototypes recompute single inference requests or, in proof-of-learning, selected training steps. The main obstacles are that workloads and network traffic are not reproducible by default, that the recomputation server must be secured, and that compute outside declared packets is not covered. The plan itself does not verify that spare compute is unused, and notes that non-compliant work might be hidden inside compliant-looking workloads.

Readinessmedium confidence

The design, its claim and its assumptions are public, but no implementation of whole-workload packets has been built.

Rubric assessment

Assessed use: recomputing whole workloads to show a cluster runs only declared inference

  • R1 met: the AI 2040 verification plan describes the design (discrete, reproducible packets visible to a recomputation server, with random partial recomputation), the claim it serves (correct outputs in an inference-only regime, later R&D verification) and its assumptions (reproducibility, an intact recomputation server, physical security) 1. Amodo gives a comparable written design at the level of single training steps 2. The plan is the only Implementation record for this mechanism (AI 2040 inference-only verification stack, assessed R1).
  • R2 not met: no public implementation or end-to-end result organizes whole workloads into reproducible packets. The nearest demonstrations recompute single inference requests 3 or, in proof-of-learning, selected training steps (Training-transcript verification (proof-of-learning)) 6. The plan's companion page lists a reproducible inference stack and network reproducibility as not started 5, and Amodo rates network reproducibility "not on track" 4.

Confidence is medium. The design is described only at a high level, but the plan's authors and Amodo both list the reproducible inference stack it needs as not started 4 5.

Gaps to the next level
  • A public implementation, or reproducible end-to-end results, of packet-based recomputation beyond single inference requests, under realistic model scale, hardware or a stated adversary.

Assessed 2026-09-25 against rubric v1.1.

On this page

How it works

For an inference-only retrofit, the AI 2040 verification plan proposes passive network taps that copy all input and output traffic to a recomputation server 1. All AI workloads, including inference and agentic tool use, would be organized into "discrete and reproducible packets, which are transparent to the recomputation server" 1. The server recomputes random samples of the packets to check that they are correct 1. Reproducibility is required because the correctness checks depend on it, and inference is not reproducible by default 1.

For later verification of R&D, the plan breaks workloads into a hierarchy of steps: full training run, training phases, gradient steps, layer forward and backward passes, GPU kernel calls and individual instructions 1. If each step can be made reproducible (the plan notes that seeded randomness is acceptable), the granularity of evidence collection maps predictably to the level of assurance 1. The same plan would also split data centres into isolated inference units by removing much of the back-end ("east-west") networking 1; see Bandwidth limits and compartmentalization.

The plan illustrates the approach for 2034 1:

  • a pool of about 33 billion H100-equivalents, its projected build-out for that year;
  • a recomputation budget of 1%;
  • packets of 100 H100-equivalent hours.

It charts, by the deployment's size, the probability of catching a rogue internal deployment within an hour, a week or a month 1. The plan states that detection improves as packet size decreases 1.

Amodo describes a related design for pre-training, in which each training step is the unit 2:

  • The prover commits a hash of sampled weight positions before learning whether the step will be audited 2.
  • For audited steps, the verifier re-runs the step and accepts if the result agrees with the committed checkpoint within a calibrated tolerance 2.
  • Because the verifier's hardware is much slower than the prover's, the audit rate follows from the slowdown: about 0.1% of steps if recomputing a step takes 1,000 times as long as the step itself 2.

Amodo reports a prototype that recomputes individual inference requests, which Sampled inference recomputation covers 3.

What it establishes

The plan argues that random partial recomputation can verify all outputs with very high probability while using a small fraction of the original compute 1. Like other recomputation schemes, it checks whether declared workloads were computed correctly, not whether other work ran outside them 2. The plan states that it does not verify that spare compute is left unused for unapproved workloads 1.

Threat model

  • Recomputation server. The plan calls its integrity "an extremely important aspect" of the retrofit 1.
  • Physical security. The plan relies on strong physical security to protect the verification measures from tampering 1.
  • Cooperation. Amodo's schemes assume the prover cooperates and lets the verifier install a recomputation server in its data centre 2.
  • Commit before audit. In Amodo's training-step design, the prover commits its evidence before it learns which steps will be audited 2.

Evidence

  • Whole-workload packets. As of September 2026, no implementation or end-to-end result of organizing whole workloads into reproducible packets has been published.
  • Inference recomputation. Amodo reports demonstrations of inference recomputation systems that work around nondeterminism 4, including its own single-request prototype 3.
  • Training-step recomputation. In proof-of-learning, a separate line of work, the verifier reproduces a subset of training updates and accepts each one within a noise threshold 6.
  • Reproducibility status. Amodo rates a reproducible inference stack as not started 4. It rates network reproducibility as not on track, because optimizations throughout the networking stack mean that network packets are not individually reproducible by default 4. It suggests schemes that do not need exact packet replication as a possible alternative 4. The plan's own companion page, updated in July 2026, lists both items as not started 5.

Limitations

  • Reproducibility cost. Making inference reproducible "may involve some cost overhead" 1. For training, writing a checkpoint at every step would cost more than 100% overhead, which Amodo's design avoids by keeping a spare replica 2.
  • Hidden work. The plan notes that a company might try to encode a non-compliant workload inside one that looks compliant 1. Work outside declared packets, including on spare compute, is not covered 1 2.
  • Tolerance. Amodo's training-step design accepts a recomputed step within a calibrated tolerance 2. In proof-of-learning, published spoofs exploit the tolerance that verification must allow for hardware noise 6.
  • Dependencies. The scheme relies on deterministic execution (Deterministic and bit-exact inference), complete traffic capture (Network taps and certifiers) and a secure recomputation server 1.

Known flaws

Published flaws, with their severity, kind and status. How flaws are rated.

  • Spare compute is outside the schemeSignificantTheoretical argumentOpen

    The plan states that it does not verify that spare compute is not used for unapproved workloads, because this seems very challenging 1. Recomputation checks the correctness of declared work, not its completeness 2.

  • Non-compliant work could be encoded inside compliant-looking packetsSignificantTheoretical argumentOpen

    The plan notes that an AI company might try to encode a non-compliant workload inside a workload that looks compliant on the surface 1.

Blockers

Sources

  1. CR. Dean (2026). Verification Plan. AI 2040. Source recordSupports: packet design, hierarchy of workload steps, recomputation budget example, assumptions and stated gaps · 2027: Concrete inference-only retrofitting proposal; Feb 2029; Early 2030: Workload Approval, Workload Verification (2034 chart)
  2. CAmodo Design (2026). Example Schemes for Verifying High-Stakes AI Agreements. Amodo Design. Source recordSupports: step-level recomputation design for pre-training; shadow replica instead of per-step checkpoint writes; commitments before audit selection; tolerance comparison; audit rate; correctness vs completeness · pre-training scheme; introduction
  3. CAmodo Design (2026). Scaling Recomputation Inference Verification. Amodo Design. Source recordSupports: single-request inference recomputation prototype · whole note
  4. CAmodo Design (2026). AI 2040 Plan A — Verification SITREP. Amodo Design. Source recordSupports: status of reproducible inference stack and network reproducibility · status items
  5. CAI Futures Project (2026). Get Involved in Verification. AI 2040. Source recordSupports: plan authors' status of a reproducible inference stack and network reproducibility, July 2026 · reproducible packets
  6. AC. Fang et al. (2023). Proof-of-Learning is Currently More Broken Than You Think. 8th IEEE European Symposium on Security and Privacy (EuroS&P 2023). Source recordSupports: proof-of-learning verifiers reproduce a subset of updates within a noise threshold; spoofs exploit the noise tolerance · §4.1; §4.2; §6.1

Search

Full search page