FLOP accounting
FLOP accounting estimates or verifies the total number of floating-point operations (FLOP) that a training run or other workload performs, usually to compare it with a threshold set by a rule 1 2.
Shavit lists total training compute among the rules a verifier might enforce, noting that it has proven to be an indicator of model capabilities 1. US Executive Order 14110 required reporting for models trained with more than 10^26 operations 2 until its revocation in January 2025 11, and a proposed international agreement sets a prohibited threshold of 10^24 FLOP and a monitored threshold of 10^22 FLOP 3. The EU AI Act presumes that a general-purpose AI model trained with more than 10^25 FLOP has high-impact capabilities, with notification of the European Commission required since 2 August 2025 4 5, and California's SB 53 defines frontier models by more than 10^26 operations of training compute 6. Proposed ways to count or cap FLOP include:
- Hardware time. Shavit converts a FLOP threshold into accelerator-days by assuming that every accelerator runs at its full rate with perfect parallelization, a conservative assumption that gives the fewest accelerator-days a run of that size could occupy 1.
- Energy. Wasil and colleagues suggest that a data centre's measured energy use could be converted into an approximate FLOP count 7.
- Telemetry. RAND lists estimating a workload's model FLOPs utilization (MFU) and physical signature, such as power, as a research problem 8, one that bears on on-chip telemetry.
- On-chip budgets. Offline licensing ties chip use to a renewable licence carrying a compute budget 9, as in hardware performance throttling and licensing.
Sastry and colleagues call compute a good high-level proxy for the risk of general-purpose frontier systems, though not necessarily of some narrow ones, and expect thresholds to need changing as algorithms and hardware improve 2. One system overview argues that coarse totals such as FLOP counts are not enough and seeks evidence about individual workloads 10.
Related
Used in
- R1Hardware-enabled guarantees (flexHEG) and guarantee processors
- R1Hardware performance throttling and licensing
- R2On-chip telemetry from timing, memory and performance counters⚠
- R2Training-transcript verification (proof-of-learning)⚠
- R1Proofs of useful work for capacity accounting
- R2Zero-knowledge proofs of training constraints
- Compute stock is at most a declared amount
- There is no undeclared relevant compute
- A training run stayed within declared limits
Sources
- BY. Shavit (2023). What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring. arXiv. Source recordSupports: total training compute as a rule and an indicator of capabilities; a threshold of H FLOPs converted to chip-days using each chip's FLOPs per day at full, perfectly parallel use · §2.1; §3.2, Table 1
- BG. Sastry et al. (2024). Computing Power and the Governance of Artificial Intelligence. arXiv. Source recordSupports: EO 14110 reporting threshold of 10^26 operations; compute as a good high-level proxy for risk of general-purpose frontier systems; thresholds need changing with algorithmic and hardware progress · thresholds; limitations
- BA. Scher et al. (2025). An International Agreement to Prevent the Premature Creation of Artificial Superintelligence. Machine Intelligence Research Institute. Source recordSupports: Strict Threshold 10^24 FLOP and Monitored Threshold 10^22 FLOP · §4
- AEuropean Parliament & Council of the European Union (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union, OJ L, 2024/1689. Source recordSupports: EU AI Act 10^25 FLOP presumption, notification, application from 2 August 2025 · Arts 51(2), 52(1), 113(b)
- AEuropean Commission (2025). Guidelines on the scope of the obligations for general-purpose AI models established by Regulation (EU) 2024/1689 (AI Act). European Commission, Communication C(2025) 5045 final. Source recordSupports: 10^25 FLOP presumption and notification; entry into application on 2 August 2025 · §2.3.1–2.3.2; landing page
- ACalifornia State Legislature (2025). California Senate Bill 53 (2025): Transparency in Frontier Artificial Intelligence Act. Statutes of 2025, Chapter 138 (Business and Professions Code §22757.10 et seq.). Source recordSupports: frontier model defined by more than 10^26 integer or floating-point operations · §22757.11(i)
- BA. R. Wasil et al. (2024). Verification methods for international AI agreements. arXiv. Source recordSupports: facility energy estimates could be converted into an approximation of FLOPs · Energy monitoring
- BM. Baker et al. (2025). Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment. RAND Corporation. Source recordSupports: estimating MFU and physical signature (e.g. power) as a research problem · Table 2, Appendix A.6
- BG. Kulp et al. (2024). Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090. RAND Corporation. Source recordSupports: offline licensing with a renewable licence carrying a compute budget · p. viii
- BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: coarse metrics like total FLOPs insufficient; per-workload evidence sought · verification goals
- AExecutive Office of the President (2025). Executive Order 14148: Initial Rescissions of Harmful Executive Orders and Actions. Federal Register, 90 FR 8237 (document 2025-01901, published 2025-01-28). Source recordSupports: revocation of EO 14110 on 20 January 2025 · Sec. 2(ggg)