Numerical nondeterminism
Numerical nondeterminism is variation in the results of the same computation on the same inputs, across repeated runs or across hardware and software setups, that comes from floating-point arithmetic rather than from intended randomness such as sampling 1 2.
Floating-point addition is not associative, so a rounded sum depends on the order in which its terms are accumulated 1 3. GPUs leave that order, their rounding strategy and their handling of subnormal numbers unspecified, and the same matrix multiplication can give different results on different GPU architectures 1. On a single machine, a Thinking Machines post finds the LLM forward pass run-to-run deterministic for a fixed batch, and traces the variation users see to kernels whose results change with batch size, which depends on server load 3. Exact replay can therefore need a record of the original hardware model, quantization, parallelism layout, kernels and batch size 4. For a verifier, this makes legitimate variation hard to tell from real problems 2. Verification designs either tolerate it or remove it:
- Sampled inference recomputation tolerates it by comparing outputs with a trusted reference that uses the same sampling seed 2, and training-transcript verification accepts a recomputed checkpoint within a small distance of the reported one 5.
- Deterministic and bit-exact inference removes it with batch-invariant kernels 3 or software emulation that predicts, bit for bit, the outputs of dense transformer blocks on four NVIDIA GPU models 6.
Related
Used in
- R2Bounding unexplained information in outputs
- R3Deterministic and bit-exact inference
- R3Model identity attestation⚠
- R1Network taps and certifiers
- R2Training-transcript verification (proof-of-learning)⚠
- R3Sampled inference recomputation
- R2Zero-knowledge proofs of training constraints
- R2Batch-invariant inference kernels (Thinking Machines)
- R2DiFR (Divergence From Reference)
- R3Verde and RepOps (Gensyn)
- R3Pearl proof-of-useful-work blockchain
- R3TOPLOC
- The declared model is the one being served
- This compute runs inference, not training
- A training run stayed within declared limits
Sources
- AE. Badash et al. (2026). Hawkeye: Reproducing GPU-Level Non-Determinism. Proceedings of Machine Learning and Systems 8 (MLSys 2026). Source recordSupports: GPU non-determinism arises from unspecified details including rounding strategy, subnormal numbers and accumulation order, since floating-point arithmetic is not associative; results differ between GPU architectures · abstract; §1
- BA. Karvonen et al. (2025). DiFR: Inference Verification Despite Nondeterminism. ICML 2026 Workshop on Technical AI Governance Research. Source recordSupports: re-running the same inference often gives different results due to benign numerical noise; comparison against a trusted reference conditioned on the same sampling seed · abstract
- CH. He & Thinking Machines Lab (2025). Defeating Nondeterminism in LLM Inference. Thinking Machines Lab: Connectionism. Source recordSupports: floating-point non-associativity; LLM forward pass run-to-run deterministic; lack of batch invariance with load-dependent batch size as the main cause of nondeterminism in LLM inference endpoints; batch-invariant kernels · sections on non-associativity, the concurrency hypothesis and batch invariance
- BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: metadata needed for bit-exact replay (hardware SKU, quantization, parallelism, kernels, batch size) · §5.2.2
- BY. Shavit (2023). What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring. arXiv. Source recordSupports: accept a recomputed checkpoint within a small distance of the reported one · §5.1
- BN. Cankaya (2026). Bit-Exact AI Inference Verification Without Performance Tradeoffs. ICML 2026 Workshop on Technical AI Governance Research. Source recordSupports: software emulation predicting every bit of transformer forward passes across NVIDIA GPU architectures, validated on dense blocks on A100, L40, L40S and H100 · abstract; evaluation