Numerical nondeterminism

Numerical nondeterminism is variation in the results of the same computation on the same inputs, across repeated runs or across hardware and software setups, that comes from floating-point arithmetic rather than from intended randomness such as sampling 1 2.

Floating-point addition is not associative, so a rounded sum depends on the order in which its terms are accumulated 1 3. GPUs leave that order, their rounding strategy and their handling of subnormal numbers unspecified, and the same matrix multiplication can give different results on different GPU architectures 1. On a single machine, a Thinking Machines post finds the LLM forward pass run-to-run deterministic for a fixed batch, and traces the variation users see to kernels whose results change with batch size, which depends on server load 3. Exact replay can therefore need a record of the original hardware model, quantization, parallelism layout, kernels and batch size 4. For a verifier, this makes legitimate variation hard to tell from real problems 2. Verification designs either tolerate it or remove it:

Related

Used in

Sources

  1. AE. Badash et al. (2026). Hawkeye: Reproducing GPU-Level Non-Determinism. Proceedings of Machine Learning and Systems 8 (MLSys 2026). Source recordSupports: GPU non-determinism arises from unspecified details including rounding strategy, subnormal numbers and accumulation order, since floating-point arithmetic is not associative; results differ between GPU architectures · abstract; §1
  2. BA. Karvonen et al. (2025). DiFR: Inference Verification Despite Nondeterminism. ICML 2026 Workshop on Technical AI Governance Research. Source recordSupports: re-running the same inference often gives different results due to benign numerical noise; comparison against a trusted reference conditioned on the same sampling seed · abstract
  3. CH. He & Thinking Machines Lab (2025). Defeating Nondeterminism in LLM Inference. Thinking Machines Lab: Connectionism. Source recordSupports: floating-point non-associativity; LLM forward pass run-to-run deterministic; lack of batch invariance with load-dependent batch size as the main cause of nondeterminism in LLM inference endpoints; batch-invariant kernels · sections on non-associativity, the concurrency hypothesis and batch invariance
  4. BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: metadata needed for bit-exact replay (hardware SKU, quantization, parallelism, kernels, batch size) · §5.2.2
  5. BY. Shavit (2023). What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring. arXiv. Source recordSupports: accept a recomputed checkpoint within a small distance of the reported one · §5.1
  6. BN. Cankaya (2026). Bit-Exact AI Inference Verification Without Performance Tradeoffs. ICML 2026 Workshop on Technical AI Governance Research. Source recordSupports: software emulation predicting every bit of transformer forward passes across NVIDIA GPU architectures, validated on dense blocks on A100, L40, L40S and H100 · abstract; evaluation

Search

Full search page