Interconnect bandwidth

Interconnect bandwidth is the rate at which accelerators, servers or clusters can exchange data over the links between them; it is one of the measurable specifications of AI accelerators, alongside operations per second and memory capacity 1.

Large-scale training links thousands of accelerators with high-bandwidth interconnect, while efficient inference can run on dozens to low hundreds of closely connected accelerators 2. Between such pods, inference needs to pass only tokens, whereas training exchanges gradients or activations; Scher and Thiergart identify this gap as the target of bandwidth limits, as in bandwidth limits and compartmentalization 2. The gap holds only while each inference replica, including any split of the model or its experts across devices, stays within one pod. Mixture-of-experts inference that spreads experts across devices uses all-to-all communication between them 3. Inside a data centre, front-end links carry token-level inputs and outputs, while the back-end fabric between accelerators carries tensors and collective operations at much higher bandwidth, is latency-sensitive, and is harder to tap 4. US Executive Order 14110 defined reportable computing clusters partly by network connections faster than 100 Gbit/s 1. It was revoked in January 2025 5. Sastry et al. note that the detectability of compute could be undermined if decentralized training, spread across many data centres or using lower-quality compute, becomes more viable 1.

Related

Used in

Sources

  1. BG. Sastry et al. (2024). Computing Power and the Governance of Artificial Intelligence. arXiv. Source recordSupports: communication bandwidth as a chip specification alongside operations per second and memory; EO cluster definition using network connections over 100 Gbit/s; decentralized training risk · § on quantifiability and detectability; limitations
  2. BA. Scher & L. Thiergart (2025). Mechanisms to Verify International Agreements About AI Development. arXiv. Source recordSupports: large-scale training links thousands of chips with high-bandwidth interconnect, efficient inference dozens to low hundreds; between pods inference needs tokens while training transfers gradients or activations; this gap is the target of bandwidth limits · Interconnect bandwidth limits
  3. AW. Cai et al. (2025). Shortcut-connected Expert Parallelism for Accelerating Mixture of Experts. ICML 2025, Proceedings of Machine Learning Research 267. Source recordSupports: expert-parallel MoE inference involves all-to-all cross-device communication · abstract
  4. BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: front-end token-level traffic vs high-bandwidth, latency-sensitive back-end fabric that is harder to tap · inference vs training
  5. AExecutive Office of the President (2025). Executive Order 14148: Initial Rescissions of Harmful Executive Orders and Actions. Federal Register, 90 FR 8237 (document 2025-01901, published 2025-01-28). Source recordSupports: revocation of Executive Order 14110 in January 2025 · §2(ggg)

Search

Full search page