Skip to content
Synapse Horizon

Insights / Private AI

Private AI

800G Ethernet vs InfiniBand: choosing an AI cluster fabric

Synapse Horizon

· 8 min read

Top-of-rack switches with dense bundles of teal and amber fibre cables linking GPU servers in a dim AI data centre aisle

The choice of 800G Ethernet vs InfiniBand shapes how an AI cluster performs, who can run it and what it costs to grow. It only matters once a workload spans more than one GPU server. Inside a single server, GPUs talk over the server’s own internal links.

This guide explains how the two fabrics differ, how to size a leaf-spine network, and how to decide. The explainer below counts the switches and links for your GPU count.

Are 800G Ethernet and InfiniBand the same speed now?

Yes, at the port. Ethernet reached 800 Gb/s with IEEE 802.3df-2024. The IEEE Standards Association says the work took about 16 months and finished four months ahead of schedule.

The standard builds 800 Gb/s from eight lanes of 100 Gb/s signalling. It covers copper cables and backplanes, and fibre reaches of 50 m and 100 m on multimode fibre. On single-mode fibre it reaches 500 m and 2 km.

InfiniBand reached the same speed with XDR. The InfiniBand Trade Association (IBTA) released the XDR specification in October 2023. It uses 200 Gb/s per lane to give 800 Gb/s per port and 1.6 Tb/s links between switches. That doubles the bandwidth of the previous generation.

Ethernet is following with faster lanes. The IEEE P802.3dj project covers 200 Gb/s signalling and rates up to 1.6 Tb/s. In 2024 it was scheduled to finish in 2026.

How many switches does an AI cluster fabric need?

AI clusters commonly use a leaf-spine design, a form of Clos network. The 2024 SIGCOMM paper discussed below describes a two-stage Clos of this kind. GPU servers connect to leaf switches. Every leaf then connects to every spine.

For a non-blocking fabric, each leaf uses half its ports for GPUs and half for uplinks. That way, every GPU can send at full speed at the same time. Al-Fares and colleagues set out this “fat-tree” design in a 2008 paper, using identical switches at every level.

Interactive estimate

AI cluster fabric explainer: leaf–spine sizing

One per GPU network port.

Two tiers reach 2,048 endpoints.

Fabric

32 leaves, 16 spines, 1,024 leaf-to-spine links.

Leaf switches

32

Spine switches

16

Leaf–spine links

1,024

Spines (16)Leaves (32) · +16 moreGPUs attach to each leaf
Every leaf connects to every spine. Up to 16 switches per tier are drawn.

RoCEv2 over 800G Ethernet

  • RDMA runs over routed Ethernet and IP (RoCEv2). 800 Gb/s Ethernet is standardised in IEEE 802.3df-2024.
  • Ethernet is lossy by default. RoCE fabrics add Priority Flow Control and congestion control (DCQCN is the common scheme), which need careful tuning.
  • Load balancing matters: large training flows collide under plain ECMP hashing, so operators add traffic engineering or multipath.
  • Uses the same tools and skills as the rest of your data-centre network. Ultra Ethernet 1.0 (2025) adds a transport built for AI.
Assumptions
  • Non-blocking (1:1) two-tier leaf–spine: each leaf uses half its ports for endpoints and half for uplinks, and each spine uses all its ports for leaves.
  • Two-tier capacity = radix² ÷ 2 endpoints. Larger clusters need a three-tier fat-tree, which supports radix³ ÷ 4 endpoints (Al-Fares et al., ACM SIGCOMM 2008).
  • Spine count is rounded up so that each leaf spreads its uplinks evenly across all spines.
  • An endpoint is one GPU network port. Counts exclude storage, management and front-end networks.
  • The same topology maths applies to Ethernet and InfiniBand. The toggle only changes the explanation.

Indicative estimate only. Contact us for an engineered proposal.

The switch port count, or radix, sets the limits. A two-tier fabric supports radix × radix ÷ 2 endpoints. That gives 2,048 GPU ports with 64-port switches and 8,192 with 128-port switches.

Beyond that you need a third tier. The 2008 paper shows that a three-tier fat-tree supports radix³ ÷ 4 hosts. That is 65,536 with 64-port switches.

Take 1,024 GPU ports on 64-port switches. You need 32 leaves, each with 32 GPUs and 32 uplinks. That makes 1,024 leaf-to-spine links across 16 spines. With 128-port switches, the same cluster needs 16 leaves and 8 spines.

Count the cables and optics as well as the switches. Each leaf-to-spine link has two ends, so 1,024 fibre links need 2,048 transceivers. The GPU-to-leaf links add their own.

Reach decides the cable type. IEEE 802.3df defines copper for the shortest runs inside a rack. Multimode fibre covers 50 m or 100 m, and single-mode fibre 500 m or 2 km. Place leaf switches close to the GPU servers so the many short links can use copper or short-reach fibre.

Should an AI fabric be oversubscribed?

Oversubscription means fewer uplinks than downlinks on each leaf. It saves switches and optics, but it caps the bandwidth GPUs can use at the same time.

The 2008 fat-tree paper defines it clearly. At 1:1, every host can talk to any other host at full line rate. At 5:1, only 20% of host bandwidth is available for some traffic patterns. The paper noted that typical data-centre designs of the time ran at 2.5:1 to 8:1 to cut cost.

That trade-off suits web and business traffic. It suits AI training less well. A 2024 SIGCOMM paper on RoCE training networks, covered below, notes that training traffic is bursty and spreads unevenly across links. That is why the explainer above assumes a non-blocking 1:1 design.

If budget forces a compromise, oversubscribe only where traffic is light. Storage and management networks are common candidates. Keep the GPU-to-GPU fabric at full bandwidth.

Should the GPU fabric be separate from the rest of the network?

In large clusters, usually yes. The 2024 SIGCOMM paper describes a dedicated “backend” network used only for distributed training. The authors say this let them evolve, operate and scale it separately from the rest of the data-centre network.

The paper’s training racks connect to two independent networks. The backend carries GPU-to-GPU traffic. The frontend handles tasks such as data ingestion, checkpointing and logging. The same split is a sensible starting point for a smaller private cluster.

How does InfiniBand keep the fabric lossless?

InfiniBand was built for high-performance computing. A 2021 survey of data-centre routing by Besta and colleagues explains how it works.

Each link uses credit-based flow control. A sender only transmits when the receiver has buffer space, so switches do not drop packets under normal load. That removes the need for slow, software-based retransmission of the kind TCP uses.

A central controller, the subnet manager, discovers the fabric and computes the routes. It also watches for failures. One subnet can hold up to 49,151 endpoints.

InfiniBand also supports RDMA natively. RDMA lets one server read or write another’s memory without involving its CPU. The XDR release adds better congestion control using network probes.

The trade-offs are practical. The survey notes that InfiniBand switches offer limited support for common Ethernet features such as VLANs and firewalling. It is also a separate fabric with its own tools, so your team needs the skills to run it.

How does RoCEv2 make Ethernet work for AI?

Traditional Ethernet is lossy: when a buffer fills, packets are dropped. For AI traffic, that hurts. RDMA over Converged Ethernet (RoCE) brings RDMA to Ethernet, and RoCEv2 runs it over routed IP networks.

To avoid drops, RoCE fabrics usually add Priority Flow Control (PFC). PFC lets a busy switch send “pause” frames upstream until its buffers clear.

A hyperscale operator described its RoCE training networks in a 2024 SIGCOMM paper. It chose RoCE for three reasons:

  • RoCE uses the standard RDMA programming model, so training software moved across easily.
  • Ethernet let it reuse its existing data-centre designs and tools.
  • The stack rests on open standards with support from many suppliers.

The paper is also candid about the work involved. Its AI zones use a two-stage Clos design with 400G links. Plain ECMP hashing balanced AI traffic poorly, because a few large flows collided on the same links. The team added traffic engineering and more queue pairs to spread the load.

Congestion control was harder still. The paper found the common scheme, DCQCN, hard to tune for AI traffic. The team moved congestion management into the collective communication library instead.

What does Ultra Ethernet add?

The Ultra Ethernet Consortium released its Specification 1.0 on 11 June 2025. It is an Ethernet-based stack for AI and HPC. It covers transport, congestion control, RDMA, the Ethernet link and physical layers, and network security.

The goal is to make much of the RoCE tuning above part of the standard. The consortium says the specification scales to millions of endpoints and avoids lock-in to one supplier. Treat it as a direction to plan for. Check which of its features your chosen switches and network cards support today.

800G Ethernet vs InfiniBand: which should you choose?

Start with the size of your cluster. Our Team and Department bundles run on a single workstation or server. Their GPUs talk inside the box, so there is no scale-out fabric to choose.

The question arises with the Enterprise bundle, a multi-node cluster with a 400G/800G fabric. Weigh these factors:

  • Skills and tools: if your team runs Ethernet today, RoCEv2 reuses the designs and tools it already knows. InfiniBand means a second fabric to learn and monitor.
  • Tuning effort: InfiniBand is lossless out of the box. RoCE needs PFC, congestion control and load-balancing settings done well.
  • Scale: both fabrics follow the same leaf-spine maths. Pick a switch radix that keeps you in two tiers for as long as possible.
  • Workload: the case studies above are about large training jobs, where many GPUs exchange data at once. Size the fabric for the jobs you will actually run.

Plan power and cooling alongside the network. Dense GPU racks often need liquid cooling, as our guide to direct-to-chip vs immersion cooling explains. To size the servers themselves, see our private LLM hardware requirements guide.

Key takeaways

  • Both fabrics reach 800 Gb/s per port: Ethernet via IEEE 802.3df-2024 and InfiniBand via XDR.
  • InfiniBand is lossless by design, with credit-based flow control and a subnet manager.
  • RoCEv2 brings RDMA to Ethernet but needs PFC, congestion control and load-balancing tuning.
  • A non-blocking two-tier fabric supports radix² ÷ 2 GPU ports: 2,048 with 64-port switches.
  • Choose on skills, tuning effort, scale and workload, not on port speed alone.

Frequently asked questions

Is InfiniBand faster than 800G Ethernet?
Not in raw port speed. Both reach 800 Gb/s per port: Ethernet through IEEE 802.3df-2024 and InfiniBand through the XDR specification. The differences lie in how each fabric handles loss, congestion and management.
What is RoCEv2?
RoCEv2 is RDMA over Converged Ethernet, version 2. It lets servers read and write each other's memory over routed Ethernet without involving the CPU. It needs lossless tuning, such as Priority Flow Control and congestion control, to perform well.
How many GPUs can a two-tier leaf-spine fabric connect?
A non-blocking two-tier fabric supports radix squared divided by two endpoints. That is 2,048 GPU ports with 64-port switches and 8,192 with 128-port switches. Larger clusters need a third tier.
Does a single GPU server need a scale-out fabric?
No. GPUs inside one server talk over the server's own links. A scale-out fabric is needed once a model or training job spans several servers, as in our Enterprise bundle.
What is Ultra Ethernet?
Ultra Ethernet is an Ethernet-based stack for AI and HPC from the Ultra Ethernet Consortium. Its Specification 1.0 was released on 11 June 2025. It covers transport, congestion control, RDMA, link and physical layers, and security.

How we research and review our articles

Request a quote

Get a fabric design for your AI cluster

Share your project size, location and timeline, and we will come back with a sourced proposal.

Related articles