Verde and RepOps (Gensyn)
Verde is a protocol from the company Gensyn for checking machine-learning jobs, such as inference, fine-tuning or training, that a client hands to untrusted compute providers.
Several providers run the same job. If their results differ, a bisection game narrows the dispute to a single operation, which a referee re-runs to decide who is right. The client gets the correct result if at least one provider is honest.
This needs identical results on different hardware, which Gensyn reports its RepOps library provides by fixing the order of floating-point operations. Gensyn reports settling markets in its Delphi app with REE, a runtime built on RepOps whose receipts let anyone re-run the inference, and states that Verde runs in its Judge service. RepOps ships as closed binaries.
The research paper reports that RepOps roughly doubled Llama-8B inference time. As of September 2026 no independent evaluation has been published.
Gensyn reports using its public runtime in production to settle markets whose answers anyone can re-run, but all evidence of that use comes from Gensyn, and no independent security evaluation of Verde or RepOps has been published.
Rubric assessment
Assessed use: showing a delegated output came from the declared model, if one provider is honest
- R1 met: the paper states the claim, that the client obtains the correct result if at least one provider is honest, together with the dispute protocol and its assumptions 1.
- R2 met: Gensyn's Reproducible Execution Environment, which runs LLM inference with RepOps and writes receipts for re-execution, is public as binaries with an MIT-licensed SDK 4. The paper measures RepOps on NVIDIA T4, RTX 3090 and A100 GPUs, including Llama-8B inference and fine-tuning on an A100, against a stated adversary: dishonest compute providers 1. The paper's evaluation reports overheads. The claim of bitwise-identical results across hardware comes from Gensyn's posts 2.
- R3 met on the provider's own account, as a production service that others can check. Gensyn reports that Delphi, its information-market app, is live on its mainnet, that hosted REE settlement is available for partner markets, and that markets settled by open models inside REE produce receipts anyone can re-run to verify the answer 5. The post, from May 2026, is the latest documented use, and it describes no dispute process. REE's current release is v0.6.0, which Gensyn does not label alpha or beta 6. Gensyn also states that Verde and RepOps are deployed in Judge, its AI evaluation service, without saying how Judge uses them 2 3. No party other than Gensyn is documented relying on Verde or REE for a verification decision.
- R4 not met: no independent audit, red-team or peer-reviewed security analysis has been published.
Confidence is low: R3 rests on Gensyn's own posts, and REE's documentation warns that receipts from one release may not re-verify under the next 6.
- An independent public security evaluation of the Verde dispute protocol and of RepOps reproducibility across hardware, which the paper asserts but does not test.
Assessed 2026-09-25 against rubric v1.1.
On this page
What it is
Verde is a protocol from Gensyn for checking machine-learning programs, such as inference, fine-tuning and training of language models, that a client delegates to untrusted compute providers 1. It adapts refereed delegation, a cryptographic technique in which a client sends the same job to several providers and a referee settles disagreements 1. RepOps (Reproducible Operators) is the library the paper introduces to make honest providers' results bitwise identical on different hardware 1. Most of the paper's authors work at Gensyn 1.
Verde relies on bit-exact inference, and for training jobs it resolves disputes over checkpoints, as in training-transcript verification. Gensyn reports that Verde and RepOps are deployed in Judge, its service for verifiable AI evaluation 2 3. It also reports settling markets in Delphi, its information-market app, with REE 5.
How it works
- The client gives the same job to two or more providers 1, and arbitration starts only if their outputs differ 2.
- A two-level bisection game finds the first step, and then the first operation in the computational graph, on which the providers disagree 1 2.
- The referee re-runs that one operation to decide which output is correct. The authors state that this cuts the referee's communication and compute by two orders of magnitude compared with re-running the whole disputed step 1.
RepOps fixes the order of floating-point operations for common operators such as matrix multiplication, so that an honest provider's result does not depend on its GPU 1 2. Gensyn's Reproducible Execution Environment (REE) packages this for LLM inference. It writes a receipt holding the model, prompt, output and the metadata needed to reproduce the run 4. Its reproducible mode uses RepOps kernels. A separate deterministic mode uses PyTorch's deterministic algorithms and reproduces results only on the same hardware 6.
Evidence
- The paper reports RepOps overheads against PyTorch on NVIDIA T4, RTX 3090 and A100 GPUs, for DistilBERT and Llama models 1. For Llama-8B on an A100 with 80 GB, RepOps added 98% to inference time and 126% to LoRA fine-tuning time 1.
- Gensyn reports that RepOps gives bitwise-identical results across hardware, and that Verde and RepOps are deployed in Judge 2. Gensyn introduced Judge in August 2025 with a reasoning task framed as a prediction market 3.
- Gensyn reports that Delphi is live on its mainnet, that markets settled by open models inside REE produce a receipt that anyone can re-run, and that hosted REE settlement is available for partner markets 5.
- REE is public as an MIT-licensed SDK with proprietary binaries 4. Its release notes add pipeline parallelism for models of up to 72B parameters 4, and its documentation gives version 0.6.0 as current, with reproducible int8 attention and mixture-of-experts kernels 6.
Limitations
- The guarantee needs at least one honest provider. The authors state that if all providers are dishonest, the referee accepts an incorrect output 1.
- The paper's RepOps supports 32-bit floating point only, and reproducibility only when each setup runs the program on a single GPU 1. Gensyn reports that REE has since added pipeline parallelism and int8 attention 4 6.
- Receipts are tied to an REE release. The documentation warns that some v0.5.0 receipts may not re-verify with v0.6.0 6.
- Gensyn states that Verde shows the output was produced by the declared model and data, not that the model itself is correct 2.
- The RepOps kernels ship only as binaries under a proprietary licence 4.
- All evidence of production use comes from Gensyn 2 3 5.
Blockers
Reproducibility costs throughput: RepOps added 98% to Llama-8B inference time on an A100 in the paper, and Gensyn reports a threefold cut in REE's reproducible-mode overhead without absolute figures.
The providers who re-run a job and the referee need the model and data, and the guarantee holds only if at least one provider is honest.
Sources
- BA. Arun et al. (2025). Verde: Verification via Refereed Delegation for Machine Learning Programs. arXiv. Source recordSupports: refereed delegation design, guarantee and its limit, bisection, single-operator re-execution, RepOps design, overheads, FP32 and single-GPU scope, author affiliations · abstract; §1; §3.2; §4 Table 2; limitations
- CO. Ersoy (2025). Verde Verification System In Production. Gensyn research blog. Source recordSupports: deployment in Judge; bitwise reproducibility across hardware; tasks covered; what the guarantee does not cover (provider-reported)
- CGensyn (2025). Introducing Judge. Gensyn news. Source recordSupports: Judge launch with a reasoning task framed as a prediction market (provider-reported)
- BGensyn (2026). gensyn-ai/ree: Gensyn Reproducible Execution Environment (GitHub repository). GitHub. Source recordSupports: public REE binaries and SDK, licences, receipts, release limited to reproducible LLM inference, pipeline parallelism up to 72B and threefold overhead cut in v0.2.0 (provider-reported) · README; patch notes
- CD. Jedamski (2026). Building Delphi: Pricing, Settlement, and Agentic Trading. Gensyn blog. Source recordSupports: Delphi live on Gensyn's mainnet; REE settlement receipts that anyone can re-run; hosted REE settlement for partner markets (provider-reported) · settlement section
- BGensyn (2026). Reproducible Execution Environment (REE) (Gensyn documentation). Gensyn documentation. Source recordSupports: REE modes, current release v0.6.0, reproducible int8 attention and MoE kernels, receipts across releases (provider-reported) · whole page