{
  "schema_version": "1.2.0",
  "rubric_version": "1.1",
  "license": "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)",
  "record": {
    "id": "I-0015",
    "slug": "gensyn-verde-repops",
    "title": "Verde and RepOps (Gensyn)",
    "aliases": [
      "Verde",
      "RepOps",
      "Reproducible Operators",
      "Gensyn Reproducible Execution Environment (REE)"
    ],
    "status": "published",
    "last_reviewed": "2026-09-25",
    "review_interval_days": 90,
    "steward": null,
    "provenance": {
      "drafted_by": "ai",
      "reviewed_by": [
        "codex-review"
      ]
    },
    "risk_flags": [],
    "flags": [
      "provider-reported"
    ],
    "one_liner": "Gensyn's system for checking delegated machine-learning jobs, which settles disagreements between providers by re-running a single operation with bitwise-reproducible operators.",
    "summary": "Verde is a protocol from the company Gensyn for checking machine-learning jobs, such as inference, fine-tuning or training, that a client hands to untrusted compute providers. Several providers run the same job. If their results differ, a bisection game narrows the dispute to a single operation, which a referee re-runs to decide who is right. The client gets the correct result if at least one provider is honest. This needs identical results on different hardware, which Gensyn reports its RepOps library provides by fixing the order of floating-point operations. Gensyn reports settling markets in its Delphi app with REE, a runtime built on RepOps whose receipts let anyone re-run the inference, and states that Verde runs in its Judge service. RepOps ships as closed binaries. The research paper reports that RepOps roughly doubled Llama-8B inference time. As of September 2026 no independent evaluation has been published.",
    "category": "cryptographic-computational",
    "secondary_categories": [],
    "verifies": [
      {
        "claim": "C-0005",
        "role": "primary",
        "note": "A client learns that a delegated inference output came from the declared model and input, if at least one provider is honest (S-1809, S-1810)."
      },
      {
        "claim": "C-0007",
        "role": "supporting",
        "note": "Also covers training and fine-tuning jobs delegated to several providers (S-1809)."
      }
    ],
    "threat_model": "adversarial",
    "adversarial_evaluation": "analysis",
    "hardware_requirement": "none",
    "prover_cooperation": "required",
    "confidentiality": "revealing",
    "depends_on": [],
    "readiness": {
      "assessment": true,
      "level": "R3",
      "scope": "showing a delegated output came from the declared model, if one provider is honest",
      "rubric_version": "1.1",
      "rationale": "Gensyn reports using its public runtime in production to settle markets whose answers anyone can re-run, but all evidence of that use comes from Gensyn, and no independent security evaluation of Verde or RepOps has been published.\n\n- **R1** met: the paper states the claim, that the client obtains the correct result if at least one provider is honest, together with the dispute protocol and its assumptions [[S-1809]].\n- **R2** met: Gensyn's Reproducible Execution Environment, which runs LLM inference with RepOps and writes receipts for re-execution, is public as binaries with an MIT-licensed SDK [[S-1812]]. The paper measures RepOps on NVIDIA T4, RTX 3090 and A100 GPUs, including Llama-8B inference and fine-tuning on an A100, against a stated adversary: dishonest compute providers [[S-1809]]. The paper's evaluation reports overheads. The claim of bitwise-identical results across hardware comes from Gensyn's posts [[S-1810]].\n- **R3** met on the provider's own account, as a production service that others can check. Gensyn reports that Delphi, its information-market app, is live on its mainnet, that hosted REE settlement is available for partner markets, and that markets settled by open models inside REE produce receipts anyone can re-run to verify the answer [[S-3022]]. The post, from May 2026, is the latest documented use, and it describes no dispute process. REE's current release is v0.6.0, which Gensyn does not label alpha or beta [[S-3023]]. Gensyn also states that Verde and RepOps are deployed in Judge, its AI evaluation service, without saying how Judge uses them [[S-1810]] [[S-1811]]. No party other than Gensyn is documented relying on Verde or REE for a verification decision.\n- **R4** not met: no independent audit, red-team or peer-reviewed security analysis has been published.\n\nConfidence is low: R3 rests on Gensyn's own posts, and REE's documentation warns that receipts from one release may not re-verify under the next [[S-3023]].",
      "evidence": [
        "S-1809",
        "S-1810",
        "S-1811",
        "S-1812",
        "S-3022",
        "S-3023"
      ],
      "next_level_gaps": [
        "An independent public security evaluation of the Verde dispute protocol and of RepOps reproducibility across hardware, which the paper asserts but does not test."
      ],
      "confidence": "low",
      "assessed_by": [
        "claude-review",
        "codex-review"
      ],
      "assessed_on": "2026-09-25",
      "status": "current",
      "dispute": null
    },
    "flaws": [],
    "blockers": [
      {
        "text": "Reproducibility costs throughput: RepOps added 98% to Llama-8B inference time on an A100 in the paper, and Gensyn reports a threefold cut in REE's reproducible-mode overhead without absolute figures.",
        "theme": "performance-compatibility",
        "blocked_by": null,
        "sources": [
          "S-1809",
          "S-1812"
        ]
      },
      {
        "text": "The providers who re-run a job and the referee need the model and data, and the guarantee holds only if at least one provider is honest.",
        "theme": "privacy-leakage",
        "blocked_by": null,
        "sources": [
          "S-1809"
        ]
      }
    ],
    "challenge_themes": [
      "performance-compatibility",
      "adversarial-validation",
      "privacy-leakage",
      "protocol-soundness"
    ],
    "organizations": [],
    "people": [],
    "sources": [
      {
        "source": "S-1809",
        "supports": "refereed delegation design, guarantee and its limit, bisection, single-operator re-execution, RepOps design, overheads, FP32 and single-GPU scope, author affiliations",
        "locator": "abstract; §1; §3.2; §4 Table 2; limitations"
      },
      {
        "source": "S-1810",
        "supports": "deployment in Judge; bitwise reproducibility across hardware; tasks covered; what the guarantee does not cover (provider-reported)"
      },
      {
        "source": "S-1811",
        "supports": "Judge launch with a reasoning task framed as a prediction market (provider-reported)"
      },
      {
        "source": "S-1812",
        "supports": "public REE binaries and SDK, licences, receipts, release limited to reproducible LLM inference, pipeline parallelism up to 72B and threefold overhead cut in v0.2.0 (provider-reported)",
        "locator": "README; patch notes"
      },
      {
        "source": "S-3022",
        "supports": "Delphi live on Gensyn's mainnet; REE settlement receipts that anyone can re-run; hosted REE settlement for partner markets (provider-reported)",
        "locator": "settlement section"
      },
      {
        "source": "S-3023",
        "supports": "REE modes, current release v0.6.0, reproducible int8 attention and MoE kernels, receipts across releases (provider-reported)",
        "locator": "whole page"
      }
    ],
    "concepts": [
      "K-0008",
      "K-0009"
    ],
    "kind": "product",
    "developer": [
      "O-0212"
    ],
    "realises": [
      "M-0002",
      "M-0006"
    ],
    "homepage": "https://www.gensyn.ai/research/verde-verification-system-in-production",
    "repo": "https://github.com/gensyn-ai/ree",
    "type": "implementation",
    "url": "https://trustbutveri.fyi/implementations/gensyn-verde-repops/",
    "source_file": "content/implementations/gensyn-verde-repops.md",
    "flags_all": [
      "provider-reported"
    ],
    "body_markdown": "## What it is\n\nVerde is a protocol from Gensyn for checking machine-learning programs, such as inference, fine-tuning and training of language models, that a client delegates to untrusted compute providers [[S-1809]]. It adapts refereed delegation, a cryptographic technique in which a client sends the same job to several providers and a referee settles disagreements [[S-1809]]. RepOps (Reproducible Operators) is the library the paper introduces to make honest providers' results bitwise identical on different hardware [[S-1809]]. Most of the paper's authors work at Gensyn [[S-1809]].\n\nVerde relies on [[M-0002|bit-exact inference]], and for training jobs it resolves disputes over checkpoints, as in [[M-0006|training-transcript verification]]. Gensyn reports that Verde and RepOps are deployed in Judge, its service for verifiable AI evaluation [[S-1810]] [[S-1811]]. It also reports settling markets in Delphi, its information-market app, with REE [[S-3022]].\n\n## How it works\n\n1. The client gives the same job to two or more providers [[S-1809]], and arbitration starts only if their outputs differ [[S-1810]].\n2. A two-level bisection game finds the first step, and then the first operation in the computational graph, on which the providers disagree [[S-1809]] [[S-1810]].\n3. The referee re-runs that one operation to decide which output is correct. The authors state that this cuts the referee's communication and compute by two orders of magnitude compared with re-running the whole disputed step [[S-1809]].\n\nRepOps fixes the order of floating-point operations for common operators such as matrix multiplication, so that an honest provider's result does not depend on its GPU [[S-1809]] [[S-1810]]. Gensyn's Reproducible Execution Environment (REE) packages this for LLM inference. It writes a receipt holding the model, prompt, output and the metadata needed to reproduce the run [[S-1812]]. Its reproducible mode uses RepOps kernels. A separate deterministic mode uses PyTorch's deterministic algorithms and reproduces results only on the same hardware [[S-3023]].\n\n## Evidence\n- The paper reports RepOps overheads against PyTorch on NVIDIA T4, RTX 3090 and A100 GPUs, for DistilBERT and Llama models [[S-1809]]. For Llama-8B on an A100 with 80 GB, RepOps added 98% to inference time and 126% to LoRA fine-tuning time [[S-1809]].\n- Gensyn reports that RepOps gives bitwise-identical results across hardware, and that Verde and RepOps are deployed in Judge [[S-1810]]. Gensyn introduced Judge in August 2025 with a reasoning task framed as a prediction market [[S-1811]].\n- Gensyn reports that Delphi is live on its mainnet, that markets settled by open models inside REE produce a receipt that anyone can re-run, and that hosted REE settlement is available for partner markets [[S-3022]].\n- REE is public as an MIT-licensed SDK with proprietary binaries [[S-1812]]. Its release notes add pipeline parallelism for models of up to 72B parameters [[S-1812]], and its documentation gives version 0.6.0 as current, with reproducible int8 attention and mixture-of-experts kernels [[S-3023]].\n\n## Limitations\n- The guarantee needs at least one honest provider. The authors state that if all providers are dishonest, the referee accepts an incorrect output [[S-1809]].\n- The paper's RepOps supports 32-bit floating point only, and reproducibility only when each setup runs the program on a single GPU [[S-1809]]. Gensyn reports that REE has since added pipeline parallelism and int8 attention [[S-1812]] [[S-3023]].\n- Receipts are tied to an REE release. The documentation warns that some v0.5.0 receipts may not re-verify with v0.6.0 [[S-3023]].\n- Gensyn states that Verde shows the output was produced by the declared model and data, not that the model itself is correct [[S-1810]].\n- The RepOps kernels ship only as binaries under a proprietary licence [[S-1812]].\n- All evidence of production use comes from Gensyn [[S-1810]] [[S-1811]] [[S-3022]].",
    "body_text": "What it is Verde is a protocol from Gensyn for checking machine-learning programs, such as inference, fine-tuning and training of language models, that a client delegates to untrusted compute providers [S-1809]. It adapts refereed delegation, a cryptographic technique in which a client sends the same job to several providers and a referee settles disagreements [S-1809]. RepOps (Reproducible Operators) is the library the paper introduces to make honest providers' results bitwise identical on different hardware [S-1809]. Most of the paper's authors work at Gensyn [S-1809]. Verde relies on bit-exact inference, and for training jobs it resolves disputes over checkpoints, as in training-transcript verification. Gensyn reports that Verde and RepOps are deployed in Judge, its service for verifiable AI evaluation [S-1810] [S-1811]. It also reports settling markets in Delphi, its information-market app, with REE [S-3022]. How it works 1. The client gives the same job to two or more providers [S-1809], and arbitration starts only if their outputs differ [S-1810]. 2. A two-level bisection game finds the first step, and then the first operation in the computational graph, on which the providers disagree [S-1809] [S-1810]. 3. The referee re-runs that one operation to decide which output is correct. The authors state that this cuts the referee's communication and compute by two orders of magnitude compared with re-running the whole disputed step [S-1809]. RepOps fixes the order of floating-point operations for common operators such as matrix multiplication, so that an honest provider's result does not depend on its GPU [S-1809] [S-1810]. Gensyn's Reproducible Execution Environment (REE) packages this for LLM inference. It writes a receipt holding the model, prompt, output and the metadata needed to reproduce the run [S-1812]. Its reproducible mode uses RepOps kernels. A separate deterministic mode uses PyTorch's deterministic algorithms and reproduces results only on the same hardware [S-3023]. Evidence - The paper reports RepOps overheads against PyTorch on NVIDIA T4, RTX 3090 and A100 GPUs, for DistilBERT and Llama models [S-1809]. For Llama-8B on an A100 with 80 GB, RepOps added 98% to inference time and 126% to LoRA fine-tuning time [S-1809]. - Gensyn reports that RepOps gives bitwise-identical results across hardware, and that Verde and RepOps are deployed in Judge [S-1810]. Gensyn introduced Judge in August 2025 with a reasoning task framed as a prediction market [S-1811]. - Gensyn reports that Delphi is live on its mainnet, that markets settled by open models inside REE produce a receipt that anyone can re-run, and that hosted REE settlement is available for partner markets [S-3022]. - REE is public as an MIT-licensed SDK with proprietary binaries [S-1812]. Its release notes add pipeline parallelism for models of up to 72B parameters [S-1812], and its documentation gives version 0.6.0 as current, with reproducible int8 attention and mixture-of-experts kernels [S-3023]. Limitations - The guarantee needs at least one honest provider. The authors state that if all providers are dishonest, the referee accepts an incorrect output [S-1809]. - The paper's RepOps supports 32-bit floating point only, and reproducibility only when each setup runs the program on a single GPU [S-1809]. Gensyn reports that REE has since added pipeline parallelism and int8 attention [S-1812] [S-3023]. - Receipts are tied to an REE release. The documentation warns that some v0.5.0 receipts may not re-verify with v0.6.0 [S-3023]. - Gensyn states that Verde shows the output was produced by the declared model and data, not that the model itself is correct [S-1810]. - The RepOps kernels ship only as binaries under a proprietary licence [S-1812]. - All evidence of production use comes from Gensyn [S-1810] [S-1811] [S-3022].",
    "referenced_by": [
      {
        "id": "M-0002",
        "title": "Deterministic and bit-exact inference",
        "url": "https://trustbutveri.fyi/mechanisms/deterministic-inference/"
      },
      {
        "id": "M-0012",
        "title": "Model identity attestation",
        "url": "https://trustbutveri.fyi/mechanisms/model-identity-attestation/"
      },
      {
        "id": "M-0006",
        "title": "Training-transcript verification (proof-of-learning)",
        "url": "https://trustbutveri.fyi/mechanisms/proof-of-learning/"
      },
      {
        "id": "I-0016",
        "title": "Batch-invariant inference kernels (Thinking Machines)",
        "url": "https://trustbutveri.fyi/implementations/batch-invariant-inference-kernels/"
      },
      {
        "id": "O-0212",
        "title": "Gensyn",
        "url": "https://trustbutveri.fyi/organizations/gensyn/"
      }
    ]
  }
}