Readiness levels

Every mechanism and implementation has a readiness level from R0 to R4. The level says how close a technique is to being usable for its verification purpose, not how mature the underlying technology is.

Each level lists the records currently at it; ⚠ marks a record with an open critical flaw.

R0Idea

Mentioned or sketched, without a written design, claim and assumptions.

No record is at this level.

R1Proposed

A design or theory is publicly described, including the claim it would verify and its assumptions.

Roughly technology readiness levels 1–3.

R2Demonstrated

A public working implementation, or reproducible published end-to-end results, obtained under conditions representative of the verification use in at least one key respect: realistic model or cluster scale, realistic hardware, or a stated adversary. Lab and research demonstrations reach this level; production use is R3.

Roughly technology readiness levels 4–6.

R3In production

R2 holds, and it is production-grade and available, or a party other than its developer relies on it for a verification decision. Critical flaws may still be open: production use does not show that it holds up against a prover who tries to cheat.

Roughly technology readiness levels 7–8.

R4Deployment-ready

R3 holds, and at least one independent public evaluation (an audit, a red-team or a peer-reviewed security analysis) left no critical flaw open.

Roughly technology readiness level 9.

No record is at this level.

How levels are assigned

  • A record gets the highest level whose criteria all hold.
  • The level is assessed for the verification use stated beside it. A commercial component does not put a verification workflow built on it in production. TEE attestation is in production (R3) because services rely on it to show users which software runs, but that does not make TEE-based verification between rival states deployment-ready (R4).
  • "Reproducible" means the method, setup and parameters are published in enough detail for an independent team to repeat the work. Public code is not required, but its absence is noted.
  • An open critical flaw does not by itself lower R1, R2 or R3, which describe development, demonstration and use. A break that invalidates the evidence a level rests on does lower it.
  • "Production-grade" (R3) covers a system its developer runs for real users, even at a 0.x version. A release its developer labels alpha, beta or preview does not count unless someone else relies on it for a verification decision.
  • Production use that has stopped still counts while the system remains available. The rationale dates the last documented use, and confidence is low if none is documented in the last 12 months.
  • A mechanism is at least as ready as its most mature implementation for that use, and its rationale names that implementation.
  • Ratings can go down.
  • Adversarial evaluation is also judged for the verification use. Attacks on related products in other settings are context, not evidence.

What each rating records

Each rating states the verification use it is assessed for, and comes with a rationale that walks through the criteria, the sources it rests on, the gaps to the next level, a confidence, who assessed it and when, and a status (current, under review or disputed). Levels are editorial judgments under rubric v1.1, kept apart from the facts, and open to correction.

Confidence is how sure the editors are of the level: 13 low confidence, 29 medium confidence and 0 high confidence. The Status page lists the low-confidence ratings. The Methodology page covers flaws, citations and neutrality.

Search

Full search page