Compartmentalization

Compartmentalization divides a facility's accelerators into groups and restricts communication between the groups, so that groups are hard to combine into a larger workload than is allowed 1 3.

Scher and Thiergart describe pods of well-connected chips. Between pods, inference needs to pass only tokens, whereas other forms of parallelism transfer activations or gradients 1. That holds only while each inference replica, including any split of the model or its experts across devices, stays within one pod. Mixture-of-experts inference that spreads experts across devices uses all-to-all communication between them 2. Efficient inference fits within dozens to low hundreds of closely connected accelerators, while large-scale training links thousands 1. RAND's "fixed set" design likewise restricts networking so that small, fixed sets of GPUs cannot be combined into large clusters 3, and Sastry and colleagues list physical limits on chip-to-chip networking as a way to enforce compute caps 4. These ideas underlie bandwidth limits and compartmentalization.

Compartments can also separate trust domains. One low-trust design air-gaps its evaluation environments and uses optical splitters and data diodes, simple components that can be inspected for tampering, to enforce one-way data movement 5. The boundaries can be checked by observing traffic between accelerators with network taps 6, while physical channels that could bypass monitored links are the target of side-channel suppression 7.

Related

Used in

Sources

  1. BA. Scher & L. Thiergart (2025). Mechanisms to Verify International Agreements About AI Development. arXiv. Source recordSupports: pods of well-connected chips; between pods, inference needs only tokens while other forms of parallelism transfer activations or gradients; efficient inference on dozens to low hundreds of chips versus thousands for large training · Interconnect bandwidth limits
  2. AW. Cai et al. (2025). Shortcut-connected Expert Parallelism for Accelerating Mixture of Experts. ICML 2025, Proceedings of Machine Learning Research 267. Source recordSupports: expert-parallel MoE inference involves all-to-all cross-device communication · abstract
  3. BG. Kulp et al. (2024). Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090. RAND Corporation. Source recordSupports: fixed-set HEM restricting networking of small, fixed sets of GPUs · p. viii
  4. BG. Sastry et al. (2024). Computing Power and the Governance of Artificial Intelligence. arXiv. Source recordSupports: compute caps enforced via physical limits on chip-to-chip networking · enforcement mechanisms
  5. BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: air-gapped evaluation environments; optical splitters and data diodes as inspectable components enforcing one-way data movement · system architecture
  6. BM. Baker et al. (2025). Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment. RAND Corporation. Source recordSupports: network taps observing data exchanged between chips · §4.2
  7. CN. Cankaya (2026). Suppressing Side Channels in an Untrusted Data Center via Retrofitted Defenses. MIRI Technical Governance Team. Source recordSupports: physical channels could bypass network monitoring · side channels of concern

Search

Full search page