Deterministic and bit-exact inference

R3In productionProvider-reported

Deterministic inference makes an AI model's outputs reproducible bit for bit, so a verifier's re-run must match exactly.

Otherwise, re-runs on the same input often differ slightly, because floating-point results depend on the order of operations, which shifts with batch size, hardware and software. This noise forces recomputation checks to accept approximate matches, which a cheating provider could exploit. Two methods remove it: kernels whose results do not depend on batch size, or recording enough about hardware and software for a verifier to reproduce every rounding error.

vLLM and SGLang offer deterministic modes, and Gensyn and Eigen Labs report verification services built on exact replay. A published emulator predicts, bit for bit, the outputs of dense transformer blocks on four NVIDIA GPU models. The obstacles are the throughput cost of batch-invariant kernels, gaps in the emulator's coverage, no independent security evaluation, and the provider's need to disclose its full configuration.

Readinesslow confidence

Gensyn reports using exact replay in production so that anyone can check markets it settles with REE, but the production evidence comes from Gensyn, and no independent security evaluation of bit-exact verification has been published.

Rubric assessment

Assessed use: exact-replay checks that the declared model and setup produced the outputs

  • R1 met: the verification claim, a covert-adversary threat model and the information a verifier needs are published 1.
  • R2 met: public code predicts dense LLM blocks bit for bit on A100, L40, L40S and H100 GPUs running unmodified vLLM and Hugging Face engines, against a stated adversary 1. Batch-invariant modes are public in vLLM (beta) and SGLang 11 12, and batch invariance was shown on a 235-billion-parameter model 8. Exact CPU reproduction of GPU matrix multiplication is peer-reviewed 9.
  • R3 met through Gensyn's REE, on its developer's account. Gensyn reports that Delphi, its information-market app, is live on its mainnet, and that markets settled by open models inside REE produce receipts that anyone can re-run to verify the answer 21. The latest documented use is from May 2026 21. REE is publicly available 18, and its current release is v0.6.0, which Gensyn does not label alpha or beta 22. Eigen Labs' service does not count towards R3, because Eigen Labs launched it as a mainnet alpha 20. No party other than a developer is documented relying on exact replay for a verification decision.
  • R4 not met: as of September 2026 no independent audit, red-team or peer-reviewed security analysis of bit-exact verification has been published.

Confidence is low: R3 rests on Gensyn's account of its own service.

Gaps to the next level
  • An independent public evaluation (audit, red-team or peer-reviewed security analysis) of bit-exact verification that leaves no critical flaw open.

Assessed 2026-09-25 against rubric v1.1.

On this page

How it works

Floating-point arithmetic is not associative, so the same sum computed in a different order can round differently 5 8. In LLM serving, the order changes with 4:

  • batch size and kernel strategy, which depend on server load;
  • GPU type, CUDA version and kernel implementations;
  • in mixture-of-experts models, routing that depends on other tokens in the batch.

Thinking Machines Lab argues that the main reason inference endpoints are nondeterministic is that load, and so batch size, varies while kernels are not batch-invariant 8. The bit-exact work separates true nondeterminism, caused by atomic functions, from non-invariance: deterministic computation that follows different reduction trees 1.

There are two routes to exact results:

  • Batch-invariant kernels fix the reduction order for each output element whatever the batch size 8. vLLM offers a batch-invariant mode, currently in beta 12, and SGLang offers a deterministic inference mode 11. DeepSeek reports end-to-end bitwise batch-invariant and deterministic kernels, built with the goal of bitwise alignment among its pre-training, post-training and inference pipelines 10. LLM-42 enforces determinism through scheduling, replaying candidate tokens under a fixed reduction schedule, and mostly reuses existing kernels unchanged 13.
  • Record and replay uses stock engines, which already give deterministic outputs that a verifier can reproduce bit for bit, if the verifier knows the key factors and no atomic functions are called 1. The factors are the hardware model, exact weights, parallelism layout, software versions and the batch size of each forward pass 1 3. Software emulation removes the need for identical hardware 1, and Hawkeye re-executes GPU matrix multiplications on a CPU without precision loss 9.

For verification, exactness turns a recomputation check (Sampled inference recomputation) into a pass/fail test 1. Zero-knowledge proofs of inference need determinism as a precondition 1, and packet-based schemes (Whole-workload recomputation (reproducible packets)) need workloads to be reproducible 6.

What it establishes

Under exact replay, the accumulated rounding errors become an auditable signature of the software and hardware used 1. The bit-exact work names three attacks that exploit the tolerance of approximate checks: steganography, unreported changes to inference software, and covert computation in unreported batch elements 1. It argues that statistical schemes can bound the covert bandwidth these leave, but cannot close it 1.

Determinism does not capture traffic or choose samples; those come from recording and sampling mechanisms such as Network taps and certifiers and Sampled inference recomputation 2 6. Batch-invariant kernels give identical results only while the model, inference implementation and device stay fixed 4.

For varied inference stacks and mixed GPU types, DiFR's authors expect statistical verification to remain necessary 4. TAO, peer-reviewed at EuroSys 2026, is a tolerance-based alternative for heterogeneous hardware 23. It accepts each operator's output if it falls within bounds derived from IEEE-754 worst-case error and empirical profiles, and settles disputes with a Merkle-anchored dispute game 23. Its authors deployed its contract layer on an Ethereum testnet 23.

Threat model

  • Adversary. The bit-exact work targets covert adversaries, who comply with monitoring only when the chance of detection is high 1.
  • Full disclosure. Exact replay assumes the verifier learns every factor that affects the numerics 1. A reference architecture for low-trust verification lists the same replay metadata 3. Recording the batch size is described as negligible overhead for the prover 1.
  • No atomic functions. Backends must avoid atomic functions 1.
  • Correct hardware model. Cross-hardware emulation assumes the hardware's rounding, subnormal handling and accumulation order have been characterized correctly 1 9.

Evidence

  • Nondeterminism measured. For 1,000 temperature-0 completions of one prompt on Qwen3-235B-A22B, Thinking Machines reports 80 unique outputs with default kernels 8. With batch-invariant kernels, all 1,000 were identical 8.
  • Bit-exact emulation. On Qwen3 4B blocks, the emulator reports zero BF16 differences for feed-forward blocks on A100, L40, L40S and H100 GPUs 1. It reports zero differences out of 71 million elements for FlashAttention-2 at 4,000 tokens 1. The paper received a best-paper award at the ICML 2026 TAIGR workshop, and its code is public 1.
  • Matrix multiplication on CPU. Hawkeye, peer-reviewed at MLSys 2026, reports 100% success replicating 4096 × 4096 matrix multiplications on Ampere, Hopper and Lovelace GPUs 9.
  • Engines. vLLM documents its batch-invariant mode 12. SGLang reports an average slowdown of 34.35% for its deterministic mode on FlashInfer and FlashAttention 3 backends 11.
  • Without determinism. In DiFR's tests with synchronized seeds, over 98% of tokens already match exactly between provider and verifier 4.
  • Batch-invariant kernels. Thinking Machines' MIT-licensed kernels underpin SGLang's deterministic mode, and vLLM's developers state that its batch-invariant mode is based on the same work 11 14 15.
  • Verde and RepOps. Gensyn fixes the order of floating-point operations so that honest compute providers get bitwise-identical results, and settles their disagreements by re-running a single operation 16. It reports running Verde and RepOps in production 17, and settling markets in its Delphi app with REE, its reproducible runtime, whose receipts anyone can re-run 21.
  • EigenAI. Eigen Labs reports a deterministic engine built on llama.cpp with its own matrix-multiplication and reduction kernels. Its outputs were bitwise identical across 10,000 runs on one GPU model, including across hosts, and never matched between A100 and H100 GPUs 19. An optimistic protocol has a committee re-execute challenged outputs inside TEEs 19. Eigen Labs launched the service on mainnet in September 2025, as an alpha whose stake was not yet exposed to slashing 20.

Limitations

  • Throughput. In Thinking Machines' test on Qwen3-8B, vLLM's default took 26 s, the unoptimized deterministic build 55 s, and the build with an improved attention kernel 42 s 8.
  • Coverage. The emulator does not yet cover mixture-of-experts inference, non-NVIDIA GPUs, the proprietary nvjet kernel family on Hopper, or training 1. Hawkeye covers matrix multiplication only. Attention and convolutions need further reverse engineering 9. vLLM's batch-invariant mode is in beta, and its tracking issue lists open work on AMD hardware, NVFP4 and speculative decoding 12 15.
  • Residual nondeterminism. Some integer de-quantization kernels use atomic additions and remain truly nondeterministic 1.
  • Disclosure. Replay requires exact weights and configuration details 1.
  • Maturity. Amodo's status page for the AI 2040 verification plan rates a reproducible inference stack, and red-teaming of recomputation schemes, as not started 7.

Known flaws

Published flaws, with their severity, kind and status. How flaws are rated.

  • Some kernels remain genuinely nondeterministicMinorOpen questionOpen

    The bit-exact work separates kernels that are deterministic but not batch-invariant from truly nondeterministic ones that use atomic functions. Some integer de-quantization kernels use atomic additions and remain nondeterministic, so exact replay needs backends that avoid them 1.

  • Cross-hardware replay relies on reverse-engineered, closed behaviourSignificantOpen questionOpen

    Emulating one GPU's rounding on another requires reverse-engineering tensor-core arithmetic and modelling proprietary kernel choices. Hawkeye covers a subset of NVIDIA architectures and states that attention and other higher-level operations need further reverse engineering 9. For the bit-exact emulator, a proprietary Hopper kernel family is an open edge case 1.

Blockers

  • Batch-invariant kernels cost throughput: in Thinking Machines' Qwen3-8B test, an improved deterministic build took 42 s against 26 s for vLLM's default, and SGLang reports an average 34.35% slowdown on its FlashInfer and FlashAttention 3 backends.

  • Coverage is incomplete: the bit-exact emulator targets dense blocks on NVIDIA GPUs and excludes mixture-of-experts inference and training, and vLLM's batch-invariant mode is in beta, with open work on AMD hardware and speculative decoding.

  • Amodo's status page for the AI 2040 verification plan rates a reproducible inference stack for that plan as 'not started'.

  • Exact replay requires the prover to disclose weights, software versions, parallelism and batch sizes to whoever recomputes.

Technical detail

Two routes lead to exact results.

  • Invariance. Kernels fix the reduction order for each output element regardless of batch size. Thinking Machines made RMSNorm, matrix multiplication and attention batch-invariant, the last with a fixed split size for the key-value dimension rather than a fixed number of splits 8. vLLM exposes this behind VLLM_BATCH_INVARIANT=1 on NVIDIA GPUs of compute capability 8.0 or higher and on Intel XPUs, in beta 12. SGLang integrated batch-invariant attention for its FlashInfer, FlashAttention 3 and Triton backends 11. LLM-42 instead decodes on a non-deterministic fast path and replays candidate tokens under a fixed-shape reduction schedule, rolling back any that are inconsistent 13.
  • Record and replay. Stock engines are deterministic but not invariant. Outputs are bitwise reproducible if the verifier knows the hardware model, the exact deployed weights, the parallelism topology (separately for prefill and decode), software versions including custom kernels, and the batch size of each forward pass 1 3. Of these, only batch size changes during serving, and it costs one extra integer per forward pass to record 1. A software emulator reproduces the rounding of other GPU models by modelling tensor-core accumulation and kernel-specific reduction trees 1. Hawkeye reproduces tensor-core matrix multiplication exactly on a CPU for Ampere, Hopper and Ada Lovelace GPUs in FP16, BF16 and FP8 9.

With exact replay, verification is pass/fail, and the chance of catching at least one false output in k samples is 1 − (1 − p)^k for a false-output rate p 1.

Sources

  1. BN. Cankaya (2026). Bit-Exact AI Inference Verification Without Performance Tradeoffs. ICML 2026 Workshop on Technical AI Governance Research. Source recordSupports: covert-adversary threat model; deterministic but non-invariant engines; required metadata; software emulator and its results; limitations; comparison with statistical schemes; best-paper award and public code (arXiv comments) · abstract; §1; results; limitations section; arXiv comments
  2. BR. Rinberg et al. (2025). Verifying LLM Inference to Detect Model Weight Exfiltration. arXiv. Source recordSupports: logging of inferences and random sampling for verification as components separate from recomputation · §4.2, assumptions 2 and 3 (v3)
  3. BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: replay metadata in a low-trust verification system · §5.2.2
  4. BA. Karvonen et al. (2025). DiFR: Inference Verification Despite Nondeterminism. ICML 2026 Workshop on Technical AI Governance Research. Source recordSupports: sources of benign nondeterminism; fixed-hardware determinism vs heterogeneous deployments; >98% token agreement · §2; §3; §7
  5. CAmodo Design (2026). Example Schemes for Verifying High-Stakes AI Agreements. Amodo Design. Source recordSupports: floating-point non-associativity; fuzzy comparison in recomputation schemes · determinism discussion
  6. CR. Dean (2026). Verification Plan. AI 2040. Source recordSupports: reproducibility required for packet correctness checks · Concrete inference-only retrofitting proposal
  7. CAmodo Design (2026). AI 2040 Plan A — Verification SITREP. Amodo Design. Source recordSupports: status of reproducible inference stack · status items
  8. CH. He & Thinking Machines Lab (2025). Defeating Nondeterminism in LLM Inference. Thinking Machines Lab: Connectionism. Source recordSupports: batch invariance as main cause; kernels made invariant; Qwen3-235B experiment; timings · whole post
  9. AE. Badash et al. (2026). Hawkeye: Reproducing GPU-Level Non-Determinism. Proceedings of Machine Learning and Systems 8 (MLSys 2026). Source recordSupports: exact CPU reproduction of tensor-core matrix multiplication; scope limits · abstract; §8; §9
  10. BDeepSeek-AI (2026). DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence. arXiv. Source recordSupports: provider-reported end-to-end batch-invariant and deterministic kernels · §3.3
  11. CThe SGLang Team (2025). Towards Deterministic Inference in SGLang and Reproducible RL Training. LMSYS Org blog. Source recordSupports: SGLang deterministic mode and its overhead · whole post
  12. BvLLM project (2026). Batch Invariance (vLLM documentation). vLLM documentation (GitHub, docs/features/batch_invariance.md). Source recordSupports: vLLM batch-invariance flag, supported hardware (NVIDIA compute capability 8.0+, Intel XPUs), beta status · whole page
  13. BR. Gond et al. (2026). LLM-42: Enabling Determinism in LLM Inference with Verified Speculation. arXiv. Source recordSupports: scheduling-based determinism alternative that mostly reuses existing kernels · abstract
  14. BThinking Machines Lab (2025). thinking-machines-lab/batch_invariant_ops (GitHub repository). GitHub. Source recordSupports: Thinking Machines' batch-invariant kernel library (MIT) · README
  15. CvLLM project contributors (2025). [Feature]: Batch Invariant Feature and Performance Optimization (vLLM issue #27433). GitHub (vllm-project/vllm issues). Source recordSupports: vLLM developers' statement that batch-invariance support is based on the Thinking Machines post; open work on AMD hardware, NVFP4 and speculative decoding
  16. BA. Arun et al. (2025). Verde: Verification via Refereed Delegation for Machine Learning Programs. arXiv. Source recordSupports: Verde refereed delegation and RepOps reproducible operators · abstract; §3.2
  17. CO. Ersoy (2025). Verde Verification System In Production. Gensyn research blog. Source recordSupports: Gensyn's report of Verde and RepOps in production (provider-reported)
  18. BGensyn (2026). gensyn-ai/ree: Gensyn Reproducible Execution Environment (GitHub repository). GitHub. Source recordSupports: REE public as binaries with an MIT-licensed SDK (provider-reported) · README
  19. BD. Ribeiro Alves et al. (2026). EigenAI: Deterministic Inference, Verifiable Results. arXiv. Source recordSupports: EigenAI deterministic engine and optimistic re-execution protocol; same-SKU and A100-vs-H100 determinism results (provider-reported) · abstract; Table 5
  20. CEigenCloud (2025). EigenCloud Brings Verifiable AI to Mass Market with EigenAI and EigenCompute Launches. Eigen Labs blog. Source recordSupports: EigenAI mainnet alpha launch; stake not yet exposed to slashing (provider-reported)
  21. CD. Jedamski (2026). Building Delphi: Pricing, Settlement, and Agentic Trading. Gensyn blog. Source recordSupports: Delphi live on Gensyn's mainnet; REE settlement receipts that anyone can re-run (provider-reported) · settlement section
  22. BGensyn (2026). Reproducible Execution Environment (REE) (Gensyn documentation). Gensyn documentation. Source recordSupports: REE current release v0.6.0, with no alpha or beta label (provider-reported) · whole page
  23. AJ. Yao et al. (2026). TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural Networks. Proceedings of the 21st European Conference on Computer Systems (EuroSys 2026), pp. 1515-1532. Source recordSupports: TAO: operator-level tolerance bounds instead of bitwise equality on heterogeneous hardware; Merkle-anchored dispute game; Ethereum testnet deployment · abstract; §1; evaluation

Search

Full search page