{
  "schema_version": "1.2.0",
  "rubric_version": "1.1",
  "license": "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)",
  "record": {
    "id": "M-0002",
    "slug": "deterministic-inference",
    "title": "Deterministic and bit-exact inference",
    "aliases": [
      "Bit-exact inference",
      "Batch-invariant inference",
      "Reproducible inference"
    ],
    "status": "published",
    "last_reviewed": "2026-09-25",
    "review_interval_days": 90,
    "steward": null,
    "provenance": {
      "drafted_by": "ai",
      "reviewed_by": [
        "codex-review"
      ]
    },
    "risk_flags": [],
    "flags": [
      "provider-reported"
    ],
    "one_liner": "Making model inference reproducible bit for bit, so that a verifier's re-run must match the provider's output exactly rather than approximately.",
    "summary": "Deterministic inference makes an AI model's outputs reproducible bit for bit, so a verifier's re-run must match exactly. Otherwise, re-runs on the same input often differ slightly, because floating-point results depend on the order of operations, which shifts with batch size, hardware and software. This noise forces recomputation checks to accept approximate matches, which a cheating provider could exploit. Two methods remove it: kernels whose results do not depend on batch size, or recording enough about hardware and software for a verifier to reproduce every rounding error. vLLM and SGLang offer deterministic modes, and Gensyn and Eigen Labs report verification services built on exact replay. A published emulator predicts, bit for bit, the outputs of dense transformer blocks on four NVIDIA GPU models. The obstacles are the throughput cost of batch-invariant kernels, gaps in the emulator's coverage, no independent security evaluation, and the provider's need to disclose its full configuration.",
    "technical": "Two routes lead to exact results.\n\n- **Invariance.** Kernels fix the reduction order for each output element regardless of batch size. Thinking Machines made RMSNorm, matrix multiplication and attention batch-invariant, the last with a fixed split size for the key-value dimension rather than a fixed number of splits [[S-1009]]. vLLM exposes this behind VLLM_BATCH_INVARIANT=1 on NVIDIA GPUs of compute capability 8.0 or higher and on Intel XPUs, in beta [[S-1013]]. SGLang integrated batch-invariant attention for its FlashInfer, FlashAttention 3 and Triton backends [[S-1012]]. LLM-42 instead decodes on a non-deterministic fast path and replays candidate tokens under a fixed-shape reduction schedule, rolling back any that are inconsistent [[S-1014]].\n- **Record and replay.** Stock engines are deterministic but not invariant. Outputs are bitwise reproducible if the verifier knows the hardware model, the exact deployed weights, the parallelism topology (separately for prefill and decode), software versions including custom kernels, and the batch size of each forward pass [[S-0018]] [[S-0020]]. Of these, only batch size changes during serving, and it costs one extra integer per forward pass to record [[S-0020]]. A software emulator reproduces the rounding of other GPU models by modelling tensor-core accumulation and kernel-specific reduction trees [[S-0020]]. Hawkeye reproduces tensor-core matrix multiplication exactly on a CPU for Ampere, Hopper and Ada Lovelace GPUs in FP16, BF16 and FP8 [[S-1010]].\n\nWith exact replay, verification is pass/fail, and the chance of catching at least one false output in k samples is 1 − (1 − p)^k for a false-output rate p [[S-0020]].",
    "category": "cryptographic-computational",
    "secondary_categories": [],
    "verifies": [
      {
        "claim": "C-0005",
        "role": "primary",
        "note": "Enables exact-match recomputation checks that the declared model, weights and software setup produced the outputs."
      },
      {
        "claim": "C-0004",
        "role": "supporting",
        "note": "Bit-exact recomputation of declared inference removes the tolerance an operator could hide other work in (S-0020)."
      },
      {
        "claim": "C-0009",
        "role": "supporting",
        "note": "Removes the tolerance margin that steganographic exfiltration could use (S-0020)."
      },
      {
        "claim": "C-0010",
        "role": "supporting",
        "note": "Unreported batch elements alter the numerics, so covert computation inside batches becomes detectable (S-0020)."
      }
    ],
    "threat_model": "adversarial",
    "adversarial_evaluation": "analysis",
    "hardware_requirement": "none",
    "prover_cooperation": "required",
    "confidentiality": "partial",
    "depends_on": [],
    "readiness": {
      "assessment": true,
      "level": "R3",
      "scope": "exact-replay checks that the declared model and setup produced the outputs",
      "rubric_version": "1.1",
      "rationale": "Gensyn reports using exact replay in production so that anyone can check markets it settles with REE, but the production evidence comes from Gensyn, and no independent security evaluation of bit-exact verification has been published.\n\n- **R1** met: the verification claim, a covert-adversary threat model and the information a verifier needs are published [[S-0020]].\n- **R2** met: public code predicts dense LLM blocks bit for bit on A100, L40, L40S and H100 GPUs running unmodified vLLM and Hugging Face engines, against a stated adversary [[S-0020]]. Batch-invariant modes are public in vLLM (beta) and SGLang [[S-1013]] [[S-1012]], and batch invariance was shown on a 235-billion-parameter model [[S-1009]]. Exact CPU reproduction of GPU matrix multiplication is peer-reviewed [[S-1010]].\n- **R3** met through [[I-0015|Gensyn's REE]], on its developer's account. Gensyn reports that Delphi, its information-market app, is live on its mainnet, and that markets settled by open models inside REE produce receipts that anyone can re-run to verify the answer [[S-3022]]. The latest documented use is from May 2026 [[S-3022]]. REE is publicly available [[S-1812]], and its current release is v0.6.0, which Gensyn does not label alpha or beta [[S-3023]]. Eigen Labs' service does not count towards R3, because Eigen Labs launched it as a mainnet alpha [[S-3021]]. No party other than a developer is documented relying on exact replay for a verification decision.\n- **R4** not met: as of September 2026 no independent audit, red-team or peer-reviewed security analysis of bit-exact verification has been published.\n\nConfidence is low: R3 rests on Gensyn's account of its own service.",
      "evidence": [
        "S-0020",
        "S-1009",
        "S-1010",
        "S-1012",
        "S-1013",
        "S-1812",
        "S-3021",
        "S-3022",
        "S-3023"
      ],
      "next_level_gaps": [
        "An independent public evaluation (audit, red-team or peer-reviewed security analysis) of bit-exact verification that leaves no critical flaw open."
      ],
      "confidence": "low",
      "assessed_by": [
        "claude-review",
        "codex-review"
      ],
      "assessed_on": "2026-09-25",
      "status": "current",
      "dispute": null
    },
    "flaws": [
      {
        "assessment": true,
        "title": "Some kernels remain genuinely nondeterministic",
        "kind": "open-question",
        "severity": "minor",
        "status": "open",
        "description": "The bit-exact work separates kernels that are deterministic but not batch-invariant from truly nondeterministic ones that use atomic functions. Some integer de-quantization kernels use atomic additions and remain nondeterministic, so exact replay needs backends that avoid them [[S-0020]].",
        "sources": [
          "S-0020"
        ],
        "response": null
      },
      {
        "assessment": true,
        "title": "Cross-hardware replay relies on reverse-engineered, closed behaviour",
        "kind": "open-question",
        "severity": "significant",
        "status": "open",
        "description": "Emulating one GPU's rounding on another requires reverse-engineering tensor-core arithmetic and modelling proprietary kernel choices. Hawkeye covers a subset of NVIDIA architectures and states that attention and other higher-level operations need further reverse engineering [[S-1010]]. For the bit-exact emulator, a proprietary Hopper kernel family is an open edge case [[S-0020]].",
        "sources": [
          "S-1010",
          "S-0020"
        ],
        "response": null
      }
    ],
    "blockers": [
      {
        "text": "Batch-invariant kernels cost throughput: in Thinking Machines' Qwen3-8B test, an improved deterministic build took 42 s against 26 s for vLLM's default, and SGLang reports an average 34.35% slowdown on its FlashInfer and FlashAttention 3 backends.",
        "theme": "performance-compatibility",
        "blocked_by": null,
        "sources": [
          "S-1009",
          "S-1012"
        ]
      },
      {
        "text": "Coverage is incomplete: the bit-exact emulator targets dense blocks on NVIDIA GPUs and excludes mixture-of-experts inference and training, and vLLM's batch-invariant mode is in beta, with open work on AMD hardware and speculative decoding.",
        "theme": "performance-compatibility",
        "blocked_by": null,
        "sources": [
          "S-0020",
          "S-1013",
          "S-1814"
        ]
      },
      {
        "text": "Amodo's status page for the AI 2040 verification plan rates a reproducible inference stack for that plan as 'not started'.",
        "theme": "performance-compatibility",
        "blocked_by": null,
        "sources": [
          "S-1008"
        ]
      },
      {
        "text": "Exact replay requires the prover to disclose weights, software versions, parallelism and batch sizes to whoever recomputes.",
        "theme": "privacy-leakage",
        "blocked_by": null,
        "sources": [
          "S-0020",
          "S-0018"
        ]
      }
    ],
    "challenge_themes": [
      "performance-compatibility",
      "protocol-soundness",
      "privacy-leakage",
      "adversarial-validation"
    ],
    "organizations": [],
    "people": [],
    "sources": [
      {
        "source": "S-0020",
        "supports": "covert-adversary threat model; deterministic but non-invariant engines; required metadata; software emulator and its results; limitations; comparison with statistical schemes; best-paper award and public code (arXiv comments)",
        "locator": "abstract; §1; results; limitations section; arXiv comments"
      },
      {
        "source": "S-0015",
        "supports": "logging of inferences and random sampling for verification as components separate from recomputation",
        "locator": "§4.2, assumptions 2 and 3 (v3)"
      },
      {
        "source": "S-0018",
        "supports": "replay metadata in a low-trust verification system",
        "locator": "§5.2.2"
      },
      {
        "source": "S-0016",
        "supports": "sources of benign nondeterminism; fixed-hardware determinism vs heterogeneous deployments; >98% token agreement",
        "locator": "§2; §3; §7"
      },
      {
        "source": "S-0017",
        "supports": "floating-point non-associativity; fuzzy comparison in recomputation schemes",
        "locator": "determinism discussion"
      },
      {
        "source": "S-0067",
        "supports": "reproducibility required for packet correctness checks",
        "locator": "Concrete inference-only retrofitting proposal"
      },
      {
        "source": "S-1008",
        "supports": "status of reproducible inference stack",
        "locator": "status items"
      },
      {
        "source": "S-1009",
        "supports": "batch invariance as main cause; kernels made invariant; Qwen3-235B experiment; timings",
        "locator": "whole post"
      },
      {
        "source": "S-1010",
        "supports": "exact CPU reproduction of tensor-core matrix multiplication; scope limits",
        "locator": "abstract; §8; §9"
      },
      {
        "source": "S-1011",
        "supports": "provider-reported end-to-end batch-invariant and deterministic kernels",
        "locator": "§3.3"
      },
      {
        "source": "S-1012",
        "supports": "SGLang deterministic mode and its overhead",
        "locator": "whole post"
      },
      {
        "source": "S-1013",
        "supports": "vLLM batch-invariance flag, supported hardware (NVIDIA compute capability 8.0+, Intel XPUs), beta status",
        "locator": "whole page"
      },
      {
        "source": "S-1014",
        "supports": "scheduling-based determinism alternative that mostly reuses existing kernels",
        "locator": "abstract"
      },
      {
        "source": "S-1813",
        "supports": "Thinking Machines' batch-invariant kernel library (MIT)",
        "locator": "README"
      },
      {
        "source": "S-1814",
        "supports": "vLLM developers' statement that batch-invariance support is based on the Thinking Machines post; open work on AMD hardware, NVFP4 and speculative decoding"
      },
      {
        "source": "S-1809",
        "supports": "Verde refereed delegation and RepOps reproducible operators",
        "locator": "abstract; §3.2"
      },
      {
        "source": "S-1810",
        "supports": "Gensyn's report of Verde and RepOps in production (provider-reported)"
      },
      {
        "source": "S-1812",
        "supports": "REE public as binaries with an MIT-licensed SDK (provider-reported)",
        "locator": "README"
      },
      {
        "source": "S-3020",
        "supports": "EigenAI deterministic engine and optimistic re-execution protocol; same-SKU and A100-vs-H100 determinism results (provider-reported)",
        "locator": "abstract; Table 5"
      },
      {
        "source": "S-3021",
        "supports": "EigenAI mainnet alpha launch; stake not yet exposed to slashing (provider-reported)"
      },
      {
        "source": "S-3022",
        "supports": "Delphi live on Gensyn's mainnet; REE settlement receipts that anyone can re-run (provider-reported)",
        "locator": "settlement section"
      },
      {
        "source": "S-3023",
        "supports": "REE current release v0.6.0, with no alpha or beta label (provider-reported)",
        "locator": "whole page"
      },
      {
        "source": "S-3500",
        "supports": "TAO: operator-level tolerance bounds instead of bitwise equality on heterogeneous hardware; Merkle-anchored dispute game; Ethereum testnet deployment",
        "locator": "abstract; §1; evaluation"
      }
    ],
    "concepts": [
      "K-0008",
      "K-0009",
      "K-0016"
    ],
    "complements": [
      "M-0003",
      "M-0004"
    ],
    "alternatives": [
      "M-0001"
    ],
    "type": "mechanism",
    "implementations": [
      {
        "id": "I-0016",
        "title": "Batch-invariant inference kernels (Thinking Machines)",
        "url": "https://trustbutveri.fyi/implementations/batch-invariant-inference-kernels/"
      },
      {
        "id": "I-0015",
        "title": "Verde and RepOps (Gensyn)",
        "url": "https://trustbutveri.fyi/implementations/gensyn-verde-repops/"
      },
      {
        "id": "I-0012",
        "title": "Low-trust AI compute verification system overview",
        "url": "https://trustbutveri.fyi/implementations/low-trust-compute-verification-system-overview/"
      }
    ],
    "url": "https://trustbutveri.fyi/mechanisms/deterministic-inference/",
    "source_file": "content/mechanisms/deterministic-inference.md",
    "flags_all": [
      "provider-reported"
    ],
    "body_markdown": "## How it works\n\nFloating-point arithmetic is not associative, so the same sum computed in a different order can round differently [[S-0017]] [[S-1009]]. In LLM serving, the order changes with [[S-0016]]:\n\n- batch size and kernel strategy, which depend on server load;\n- GPU type, CUDA version and kernel implementations;\n- in mixture-of-experts models, routing that depends on other tokens in the batch.\n\nThinking Machines Lab argues that the main reason inference endpoints are nondeterministic is that load, and so batch size, varies while kernels are not batch-invariant [[S-1009]]. The bit-exact work separates true nondeterminism, caused by atomic functions, from non-invariance: deterministic computation that follows different reduction trees [[S-0020]].\n\nThere are two routes to exact results:\n\n- Batch-invariant kernels fix the reduction order for each output element whatever the batch size [[S-1009]]. vLLM offers a batch-invariant mode, currently in beta [[S-1013]], and SGLang offers a deterministic inference mode [[S-1012]]. DeepSeek reports end-to-end bitwise batch-invariant and deterministic kernels, built with the goal of bitwise alignment among its pre-training, post-training and inference pipelines [[S-1011]]. LLM-42 enforces determinism through scheduling, replaying candidate tokens under a fixed reduction schedule, and mostly reuses existing kernels unchanged [[S-1014]].\n- Record and replay uses stock engines, which already give deterministic outputs that a verifier can reproduce bit for bit, if the verifier knows the key factors and no atomic functions are called [[S-0020]]. The factors are the hardware model, exact weights, parallelism layout, software versions and the batch size of each forward pass [[S-0018]] [[S-0020]]. Software emulation removes the need for identical hardware [[S-0020]], and Hawkeye re-executes GPU matrix multiplications on a CPU without precision loss [[S-1010]].\n\nFor verification, exactness turns a recomputation check ([[M-0001]]) into a pass/fail test [[S-0020]]. [[M-0004|Zero-knowledge proofs of inference]] need determinism as a precondition [[S-0020]], and packet-based schemes ([[M-0003]]) need workloads to be reproducible [[S-0067]].\n\n## What it establishes\nUnder exact replay, the accumulated rounding errors become an auditable signature of the software and hardware used [[S-0020]]. The bit-exact work names three attacks that exploit the tolerance of approximate checks: steganography, unreported changes to inference software, and covert computation in unreported batch elements [[S-0020]]. It argues that statistical schemes can bound the covert bandwidth these leave, but cannot close it [[S-0020]].\n\nDeterminism does not capture traffic or choose samples; those come from recording and sampling mechanisms such as [[M-0013]] and [[M-0001]] [[S-0015]] [[S-0067]]. Batch-invariant kernels give identical results only while the model, inference implementation and device stay fixed [[S-0016]].\n\nFor varied inference stacks and mixed GPU types, DiFR's authors expect statistical verification to remain necessary [[S-0016]]. TAO, peer-reviewed at EuroSys 2026, is a tolerance-based alternative for heterogeneous hardware [[S-3500]]. It accepts each operator's output if it falls within bounds derived from IEEE-754 worst-case error and empirical profiles, and settles disputes with a Merkle-anchored dispute game [[S-3500]]. Its authors deployed its contract layer on an Ethereum testnet [[S-3500]].\n\n## Threat model\n- **Adversary.** The bit-exact work targets covert adversaries, who comply with monitoring only when the chance of detection is high [[S-0020]].\n- **Full disclosure.** Exact replay assumes the verifier learns every factor that affects the numerics [[S-0020]]. A reference architecture for low-trust verification lists the same replay metadata [[S-0018]]. Recording the batch size is described as negligible overhead for the prover [[S-0020]].\n- **No atomic functions.** Backends must avoid atomic functions [[S-0020]].\n- **Correct hardware model.** Cross-hardware emulation assumes the hardware's rounding, subnormal handling and accumulation order have been characterized correctly [[S-1010]] [[S-0020]].\n\n## Evidence\n- **Nondeterminism measured.** For 1,000 temperature-0 completions of one prompt on Qwen3-235B-A22B, Thinking Machines reports 80 unique outputs with default kernels [[S-1009]]. With batch-invariant kernels, all 1,000 were identical [[S-1009]].\n- **Bit-exact emulation.** On Qwen3 4B blocks, the emulator reports zero BF16 differences for feed-forward blocks on A100, L40, L40S and H100 GPUs [[S-0020]]. It reports zero differences out of 71 million elements for FlashAttention-2 at 4,000 tokens [[S-0020]]. The paper received a best-paper award at the ICML 2026 TAIGR workshop, and its code is public [[S-0020]].\n- **Matrix multiplication on CPU.** Hawkeye, peer-reviewed at MLSys 2026, reports 100% success replicating 4096 × 4096 matrix multiplications on Ampere, Hopper and Lovelace GPUs [[S-1010]].\n- **Engines.** vLLM documents its batch-invariant mode [[S-1013]]. SGLang reports an average slowdown of 34.35% for its deterministic mode on FlashInfer and FlashAttention 3 backends [[S-1012]].\n- **Without determinism.** In DiFR's tests with synchronized seeds, over 98% of tokens already match exactly between provider and verifier [[S-0016]].\n- **[[I-0016|Batch-invariant kernels]].** Thinking Machines' MIT-licensed kernels underpin SGLang's deterministic mode, and vLLM's developers state that its batch-invariant mode is based on the same work [[S-1813]] [[S-1012]] [[S-1814]].\n- **[[I-0015|Verde and RepOps]].** Gensyn fixes the order of floating-point operations so that honest compute providers get bitwise-identical results, and settles their disagreements by re-running a single operation [[S-1809]]. It reports running Verde and RepOps in production [[S-1810]], and settling markets in its Delphi app with REE, its reproducible runtime, whose receipts anyone can re-run [[S-3022]].\n- **EigenAI.** Eigen Labs reports a deterministic engine built on llama.cpp with its own matrix-multiplication and reduction kernels. Its outputs were bitwise identical across 10,000 runs on one GPU model, including across hosts, and never matched between A100 and H100 GPUs [[S-3020]]. An optimistic protocol has a committee re-execute challenged outputs inside TEEs [[S-3020]]. Eigen Labs launched the service on mainnet in September 2025, as an alpha whose stake was not yet exposed to slashing [[S-3021]].\n\n## Limitations\n- **Throughput.** In Thinking Machines' test on Qwen3-8B, vLLM's default took 26 s, the unoptimized deterministic build 55 s, and the build with an improved attention kernel 42 s [[S-1009]].\n- **Coverage.** The emulator does not yet cover mixture-of-experts inference, non-NVIDIA GPUs, the proprietary nvjet kernel family on Hopper, or training [[S-0020]]. Hawkeye covers matrix multiplication only. Attention and convolutions need further reverse engineering [[S-1010]]. vLLM's batch-invariant mode is in beta, and its tracking issue lists open work on AMD hardware, NVFP4 and speculative decoding [[S-1013]] [[S-1814]].\n- **Residual nondeterminism.** Some integer de-quantization kernels use atomic additions and remain truly nondeterministic [[S-0020]].\n- **Disclosure.** Replay requires exact weights and configuration details [[S-0020]].\n- **Maturity.** Amodo's status page for the AI 2040 verification plan rates a reproducible inference stack, and red-teaming of recomputation schemes, as not started [[S-1008]].",
    "body_text": "How it works Floating-point arithmetic is not associative, so the same sum computed in a different order can round differently [S-0017] [S-1009]. In LLM serving, the order changes with [S-0016]: - batch size and kernel strategy, which depend on server load; - GPU type, CUDA version and kernel implementations; - in mixture-of-experts models, routing that depends on other tokens in the batch. Thinking Machines Lab argues that the main reason inference endpoints are nondeterministic is that load, and so batch size, varies while kernels are not batch-invariant [S-1009]. The bit-exact work separates true nondeterminism, caused by atomic functions, from non-invariance: deterministic computation that follows different reduction trees [S-0020]. There are two routes to exact results: - Batch-invariant kernels fix the reduction order for each output element whatever the batch size [S-1009]. vLLM offers a batch-invariant mode, currently in beta [S-1013], and SGLang offers a deterministic inference mode [S-1012]. DeepSeek reports end-to-end bitwise batch-invariant and deterministic kernels, built with the goal of bitwise alignment among its pre-training, post-training and inference pipelines [S-1011]. LLM-42 enforces determinism through scheduling, replaying candidate tokens under a fixed reduction schedule, and mostly reuses existing kernels unchanged [S-1014]. - Record and replay uses stock engines, which already give deterministic outputs that a verifier can reproduce bit for bit, if the verifier knows the key factors and no atomic functions are called [S-0020]. The factors are the hardware model, exact weights, parallelism layout, software versions and the batch size of each forward pass [S-0018] [S-0020]. Software emulation removes the need for identical hardware [S-0020], and Hawkeye re-executes GPU matrix multiplications on a CPU without precision loss [S-1010]. For verification, exactness turns a recomputation check (Sampled inference recomputation) into a pass/fail test [S-0020]. Zero-knowledge proofs of inference need determinism as a precondition [S-0020], and packet-based schemes (Whole-workload recomputation (reproducible packets)) need workloads to be reproducible [S-0067]. What it establishes Under exact replay, the accumulated rounding errors become an auditable signature of the software and hardware used [S-0020]. The bit-exact work names three attacks that exploit the tolerance of approximate checks: steganography, unreported changes to inference software, and covert computation in unreported batch elements [S-0020]. It argues that statistical schemes can bound the covert bandwidth these leave, but cannot close it [S-0020]. Determinism does not capture traffic or choose samples; those come from recording and sampling mechanisms such as Network taps and certifiers and Sampled inference recomputation [S-0015] [S-0067]. Batch-invariant kernels give identical results only while the model, inference implementation and device stay fixed [S-0016]. For varied inference stacks and mixed GPU types, DiFR's authors expect statistical verification to remain necessary [S-0016]. TAO, peer-reviewed at EuroSys 2026, is a tolerance-based alternative for heterogeneous hardware [S-3500]. It accepts each operator's output if it falls within bounds derived from IEEE-754 worst-case error and empirical profiles, and settles disputes with a Merkle-anchored dispute game [S-3500]. Its authors deployed its contract layer on an Ethereum testnet [S-3500]. Threat model - Adversary. The bit-exact work targets covert adversaries, who comply with monitoring only when the chance of detection is high [S-0020]. - Full disclosure. Exact replay assumes the verifier learns every factor that affects the numerics [S-0020]. A reference architecture for low-trust verification lists the same replay metadata [S-0018]. Recording the batch size is described as negligible overhead for the prover [S-0020]. - No atomic functions. Backends must avoid atomic functions [S-0020]. - Correct hardware model. Cross-hardware emulation assumes the hardware's rounding, subnormal handling and accumulation order have been characterized correctly [S-1010] [S-0020]. Evidence - Nondeterminism measured. For 1,000 temperature-0 completions of one prompt on Qwen3-235B-A22B, Thinking Machines reports 80 unique outputs with default kernels [S-1009]. With batch-invariant kernels, all 1,000 were identical [S-1009]. - Bit-exact emulation. On Qwen3 4B blocks, the emulator reports zero BF16 differences for feed-forward blocks on A100, L40, L40S and H100 GPUs [S-0020]. It reports zero differences out of 71 million elements for FlashAttention-2 at 4,000 tokens [S-0020]. The paper received a best-paper award at the ICML 2026 TAIGR workshop, and its code is public [S-0020]. - Matrix multiplication on CPU. Hawkeye, peer-reviewed at MLSys 2026, reports 100% success replicating 4096 × 4096 matrix multiplications on Ampere, Hopper and Lovelace GPUs [S-1010]. - Engines. vLLM documents its batch-invariant mode [S-1013]. SGLang reports an average slowdown of 34.35% for its deterministic mode on FlashInfer and FlashAttention 3 backends [S-1012]. - Without determinism. In DiFR's tests with synchronized seeds, over 98% of tokens already match exactly between provider and verifier [S-0016]. - Batch-invariant kernels. Thinking Machines' MIT-licensed kernels underpin SGLang's deterministic mode, and vLLM's developers state that its batch-invariant mode is based on the same work [S-1813] [S-1012] [S-1814]. - Verde and RepOps. Gensyn fixes the order of floating-point operations so that honest compute providers get bitwise-identical results, and settles their disagreements by re-running a single operation [S-1809]. It reports running Verde and RepOps in production [S-1810], and settling markets in its Delphi app with REE, its reproducible runtime, whose receipts anyone can re-run [S-3022]. - EigenAI. Eigen Labs reports a deterministic engine built on llama.cpp with its own matrix-multiplication and reduction kernels. Its outputs were bitwise identical across 10,000 runs on one GPU model, including across hosts, and never matched between A100 and H100 GPUs [S-3020]. An optimistic protocol has a committee re-execute challenged outputs inside TEEs [S-3020]. Eigen Labs launched the service on mainnet in September 2025, as an alpha whose stake was not yet exposed to slashing [S-3021]. Limitations - Throughput. In Thinking Machines' test on Qwen3-8B, vLLM's default took 26 s, the unoptimized deterministic build 55 s, and the build with an improved attention kernel 42 s [S-1009]. - Coverage. The emulator does not yet cover mixture-of-experts inference, non-NVIDIA GPUs, the proprietary nvjet kernel family on Hopper, or training [S-0020]. Hawkeye covers matrix multiplication only. Attention and convolutions need further reverse engineering [S-1010]. vLLM's batch-invariant mode is in beta, and its tracking issue lists open work on AMD hardware, NVFP4 and speculative decoding [S-1013] [S-1814]. - Residual nondeterminism. Some integer de-quantization kernels use atomic additions and remain truly nondeterministic [S-0020]. - Disclosure. Replay requires exact weights and configuration details [S-0020]. - Maturity. Amodo's status page for the AI 2040 verification plan rates a reproducible inference stack, and red-teaming of recomputation schemes, as not started [S-1008].",
    "referenced_by": [
      {
        "id": "M-0024",
        "title": "Bounding unexplained information in outputs",
        "url": "https://trustbutveri.fyi/mechanisms/bounding-unexplained-information/"
      },
      {
        "id": "M-0012",
        "title": "Model identity attestation",
        "url": "https://trustbutveri.fyi/mechanisms/model-identity-attestation/"
      },
      {
        "id": "M-0013",
        "title": "Network taps and certifiers",
        "url": "https://trustbutveri.fyi/mechanisms/network-taps-and-certifiers/"
      },
      {
        "id": "M-0003",
        "title": "Whole-workload recomputation (reproducible packets)",
        "url": "https://trustbutveri.fyi/mechanisms/reproducible-computation-packets/"
      },
      {
        "id": "M-0001",
        "title": "Sampled inference recomputation",
        "url": "https://trustbutveri.fyi/mechanisms/sampled-inference-recomputation/"
      },
      {
        "id": "I-0016",
        "title": "Batch-invariant inference kernels (Thinking Machines)",
        "url": "https://trustbutveri.fyi/implementations/batch-invariant-inference-kernels/"
      },
      {
        "id": "I-0015",
        "title": "Verde and RepOps (Gensyn)",
        "url": "https://trustbutveri.fyi/implementations/gensyn-verde-repops/"
      },
      {
        "id": "I-0012",
        "title": "Low-trust AI compute verification system overview",
        "url": "https://trustbutveri.fyi/implementations/low-trust-compute-verification-system-overview/"
      },
      {
        "id": "I-0004",
        "title": "Pearl proof-of-useful-work blockchain",
        "url": "https://trustbutveri.fyi/implementations/pearl-proof-of-useful-work/"
      },
      {
        "id": "I-0008",
        "title": "SASH confidential network logger",
        "url": "https://trustbutveri.fyi/implementations/sash-confidential-network-logger/"
      },
      {
        "id": "C-0009",
        "title": "Model weights have not left the facility",
        "url": "https://trustbutveri.fyi/claims/weights-have-not-left/"
      },
      {
        "id": "K-0008",
        "title": "Numerical nondeterminism",
        "url": "https://trustbutveri.fyi/concepts/numerical-nondeterminism/"
      },
      {
        "id": "O-0212",
        "title": "Gensyn",
        "url": "https://trustbutveri.fyi/organizations/gensyn/"
      },
      {
        "id": "O-0202",
        "title": "Machine Intelligence Research Institute",
        "url": "https://trustbutveri.fyi/organizations/machine-intelligence-research-institute/"
      },
      {
        "id": "O-0121",
        "title": "Pearl Research Labs",
        "url": "https://trustbutveri.fyi/organizations/pearl-research/"
      },
      {
        "id": "O-0214",
        "title": "Thinking Machines Lab",
        "url": "https://trustbutveri.fyi/organizations/thinking-machines-lab/"
      }
    ]
  }
}