Insights / Private AI
Private AI800G Ethernet vs InfiniBand: choosing an AI cluster fabric
· 8 min read

On this page
The choice of 800G Ethernet vs InfiniBand shapes how an AI cluster performs, who can run it and what it costs to grow. It only matters once a workload spans more than one GPU server. Inside a single server, GPUs talk over the server’s own internal links.
This guide explains how the two fabrics differ, how to size a leaf-spine network, and how to decide. The explainer below counts the switches and links for your GPU count.
Are 800G Ethernet and InfiniBand the same speed now?
Yes, at the port. Ethernet reached 800 Gb/s with IEEE 802.3df-2024. The IEEE Standards Association says the work took about 16 months and finished four months ahead of schedule.
The standard builds 800 Gb/s from eight lanes of 100 Gb/s signalling. It covers copper cables and backplanes, and fibre reaches of 50 m and 100 m on multimode fibre. On single-mode fibre it reaches 500 m and 2 km.
InfiniBand reached the same speed with XDR. The InfiniBand Trade Association (IBTA) released the XDR specification in October 2023. It uses 200 Gb/s per lane to give 800 Gb/s per port and 1.6 Tb/s links between switches. That doubles the bandwidth of the previous generation.
Ethernet is following with faster lanes. The IEEE P802.3dj project covers 200 Gb/s signalling and rates up to 1.6 Tb/s. In 2024 it was scheduled to finish in 2026.
How many switches does an AI cluster fabric need?
AI clusters commonly use a leaf-spine design, a form of Clos network. The 2024 SIGCOMM paper discussed below describes a two-stage Clos of this kind. GPU servers connect to leaf switches. Every leaf then connects to every spine.
For a non-blocking fabric, each leaf uses half its ports for GPUs and half for uplinks. That way, every GPU can send at full speed at the same time. Al-Fares and colleagues set out this “fat-tree” design in a 2008 paper, using identical switches at every level.
Interactive estimate
AI cluster fabric explainer: leaf–spine sizing
One per GPU network port.
Two tiers reach 2,048 endpoints.
32 leaves, 16 spines, 1,024 leaf-to-spine links.
Leaf switches
32
Spine switches
16
Leaf–spine links
1,024
RoCEv2 over 800G Ethernet
- RDMA runs over routed Ethernet and IP (RoCEv2). 800 Gb/s Ethernet is standardised in IEEE 802.3df-2024.
- Ethernet is lossy by default. RoCE fabrics add Priority Flow Control and congestion control (DCQCN is the common scheme), which need careful tuning.
- Load balancing matters: large training flows collide under plain ECMP hashing, so operators add traffic engineering or multipath.
- Uses the same tools and skills as the rest of your data-centre network. Ultra Ethernet 1.0 (2025) adds a transport built for AI.
Assumptions
- Non-blocking (1:1) two-tier leaf–spine: each leaf uses half its ports for endpoints and half for uplinks, and each spine uses all its ports for leaves.
- Two-tier capacity = radix² ÷ 2 endpoints. Larger clusters need a three-tier fat-tree, which supports radix³ ÷ 4 endpoints (Al-Fares et al., ACM SIGCOMM 2008).
- Spine count is rounded up so that each leaf spreads its uplinks evenly across all spines.
- An endpoint is one GPU network port. Counts exclude storage, management and front-end networks.
- The same topology maths applies to Ethernet and InfiniBand. The toggle only changes the explanation.
Indicative estimate only. Contact us for an engineered proposal.
The switch port count, or radix, sets the limits. A two-tier fabric supports radix × radix ÷ 2 endpoints. That gives 2,048 GPU ports with 64-port switches and 8,192 with 128-port switches.
Beyond that you need a third tier. The 2008 paper shows that a three-tier fat-tree supports radix³ ÷ 4 hosts. That is 65,536 with 64-port switches.
Take 1,024 GPU ports on 64-port switches. You need 32 leaves, each with 32 GPUs and 32 uplinks. That makes 1,024 leaf-to-spine links across 16 spines. With 128-port switches, the same cluster needs 16 leaves and 8 spines.
Count the cables and optics as well as the switches. Each leaf-to-spine link has two ends, so 1,024 fibre links need 2,048 transceivers. The GPU-to-leaf links add their own.
Reach decides the cable type. IEEE 802.3df defines copper for the shortest runs inside a rack. Multimode fibre covers 50 m or 100 m, and single-mode fibre 500 m or 2 km. Place leaf switches close to the GPU servers so the many short links can use copper or short-reach fibre.
Should an AI fabric be oversubscribed?
Oversubscription means fewer uplinks than downlinks on each leaf. It saves switches and optics, but it caps the bandwidth GPUs can use at the same time.
The 2008 fat-tree paper defines it clearly. At 1:1, every host can talk to any other host at full line rate. At 5:1, only 20% of host bandwidth is available for some traffic patterns. The paper noted that typical data-centre designs of the time ran at 2.5:1 to 8:1 to cut cost.
That trade-off suits web and business traffic. It suits AI training less well. A 2024 SIGCOMM paper on RoCE training networks, covered below, notes that training traffic is bursty and spreads unevenly across links. That is why the explainer above assumes a non-blocking 1:1 design.
If budget forces a compromise, oversubscribe only where traffic is light. Storage and management networks are common candidates. Keep the GPU-to-GPU fabric at full bandwidth.
Should the GPU fabric be separate from the rest of the network?
In large clusters, usually yes. The 2024 SIGCOMM paper describes a dedicated “backend” network used only for distributed training. The authors say this let them evolve, operate and scale it separately from the rest of the data-centre network.
The paper’s training racks connect to two independent networks. The backend carries GPU-to-GPU traffic. The frontend handles tasks such as data ingestion, checkpointing and logging. The same split is a sensible starting point for a smaller private cluster.
How does InfiniBand keep the fabric lossless?
InfiniBand was built for high-performance computing. A 2021 survey of data-centre routing by Besta and colleagues explains how it works.
Each link uses credit-based flow control. A sender only transmits when the receiver has buffer space, so switches do not drop packets under normal load. That removes the need for slow, software-based retransmission of the kind TCP uses.
A central controller, the subnet manager, discovers the fabric and computes the routes. It also watches for failures. One subnet can hold up to 49,151 endpoints.
InfiniBand also supports RDMA natively. RDMA lets one server read or write another’s memory without involving its CPU. The XDR release adds better congestion control using network probes.
The trade-offs are practical. The survey notes that InfiniBand switches offer limited support for common Ethernet features such as VLANs and firewalling. It is also a separate fabric with its own tools, so your team needs the skills to run it.
How does RoCEv2 make Ethernet work for AI?
Traditional Ethernet is lossy: when a buffer fills, packets are dropped. For AI traffic, that hurts. RDMA over Converged Ethernet (RoCE) brings RDMA to Ethernet, and RoCEv2 runs it over routed IP networks.
To avoid drops, RoCE fabrics usually add Priority Flow Control (PFC). PFC lets a busy switch send “pause” frames upstream until its buffers clear.
A hyperscale operator described its RoCE training networks in a 2024 SIGCOMM paper. It chose RoCE for three reasons:
- RoCE uses the standard RDMA programming model, so training software moved across easily.
- Ethernet let it reuse its existing data-centre designs and tools.
- The stack rests on open standards with support from many suppliers.
The paper is also candid about the work involved. Its AI zones use a two-stage Clos design with 400G links. Plain ECMP hashing balanced AI traffic poorly, because a few large flows collided on the same links. The team added traffic engineering and more queue pairs to spread the load.
Congestion control was harder still. The paper found the common scheme, DCQCN, hard to tune for AI traffic. The team moved congestion management into the collective communication library instead.
What does Ultra Ethernet add?
The Ultra Ethernet Consortium released its Specification 1.0 on 11 June 2025. It is an Ethernet-based stack for AI and HPC. It covers transport, congestion control, RDMA, the Ethernet link and physical layers, and network security.
The goal is to make much of the RoCE tuning above part of the standard. The consortium says the specification scales to millions of endpoints and avoids lock-in to one supplier. Treat it as a direction to plan for. Check which of its features your chosen switches and network cards support today.
800G Ethernet vs InfiniBand: which should you choose?
Start with the size of your cluster. Our Team and Department bundles run on a single workstation or server. Their GPUs talk inside the box, so there is no scale-out fabric to choose.
The question arises with the Enterprise bundle, a multi-node cluster with a 400G/800G fabric. Weigh these factors:
- Skills and tools: if your team runs Ethernet today, RoCEv2 reuses the designs and tools it already knows. InfiniBand means a second fabric to learn and monitor.
- Tuning effort: InfiniBand is lossless out of the box. RoCE needs PFC, congestion control and load-balancing settings done well.
- Scale: both fabrics follow the same leaf-spine maths. Pick a switch radix that keeps you in two tiers for as long as possible.
- Workload: the case studies above are about large training jobs, where many GPUs exchange data at once. Size the fabric for the jobs you will actually run.
Plan power and cooling alongside the network. Dense GPU racks often need liquid cooling, as our guide to direct-to-chip vs immersion cooling explains. To size the servers themselves, see our private LLM hardware requirements guide.
Key takeaways
- Both fabrics reach 800 Gb/s per port: Ethernet via IEEE 802.3df-2024 and InfiniBand via XDR.
- InfiniBand is lossless by design, with credit-based flow control and a subnet manager.
- RoCEv2 brings RDMA to Ethernet but needs PFC, congestion control and load-balancing tuning.
- A non-blocking two-tier fabric supports radix² ÷ 2 GPU ports: 2,048 with 64-port switches.
- Choose on skills, tuning effort, scale and workload, not on port speed alone.
Frequently asked questions
Is InfiniBand faster than 800G Ethernet?
What is RoCEv2?
How many GPUs can a two-tier leaf-spine fabric connect?
Does a single GPU server need a scale-out fabric?
What is Ultra Ethernet?
Sources
- IEEE Standards Association –Ethernet's Next Bar is Now (IEEE 802.3df-2024, April 2024)
- InfiniBand Trade Association –IBTA Unveils XDR InfiniBand Specification to Enable the Next Generation of AI and Scientific Computing (October 2023)
- Ultra Ethernet Consortium –Ultra Ethernet Consortium (UEC) Launches Specification 1.0 (June 2025)
- ACM SIGCOMM (Gangidi et al.) –RDMA over Ethernet for Distributed Training at Meta Scale (ACM SIGCOMM 2024)
- UC San Diego (Al-Fares, Loukissas, Vahdat) –A Scalable, Commodity Data Center Network Architecture (ACM SIGCOMM 2008)
- ETH Zurich / RIKEN (Besta et al.), arXiv –High-Performance Routing with Multipathing and Path Diversity in Ethernet and HPC Networks (IEEE TPDS, 2021)


