News icon

Kimi K3 is now available on Runpod

NVLink vs InfiniBand vs Ethernet: GPU Interconnects Explained

Once a training job outgrows a single GPU, the interconnect stops being a footnote and starts deciding your throughput. Eight GPUs on a slow path can deliver the effective performance of four, and no amount of extra compute fixes that.

If you read one paragraph, read this one: NVLink and InfiniBand are not competitors. They operate at different layers and solve different problems. The real choice is between InfiniBand and high-speed Ethernet, and only for the layer above a single machine.

The three layers, and what belongs to each

This is the part most comparisons skip, and it makes everything after it straightforward.

LayerWhat it connectsTechnologyRough bandwidth
Inside one GPUCompute to its own memoryHBM or GDDR864 GB/s on an L40S, around 8 TB/s on a B200
Inside one serverGPU to GPU, same machineNVLink and NVSwitch, otherwise PCIeNVLink on Hopper: up to 900 GB/s per GPU. PCIe Gen4 x16: 64 GB/s bidirectional
Between serversNode to node in a clusterInfiniBand or Ethernet100 to 400 Gb/s per link is common; Runpod Clusters cite 1600 to 3200 Gbps between nodes

Note the units, because they cause more confusion than anything else here. NVLink is quoted in gigabytes per second. InfiniBand and Ethernet are quoted in gigabits per second. A factor of eight separates them before you compare anything. NVLink at 900 GB/s is roughly 7,200 Gb/s, which is why intra-node communication is not the bottleneck when a fast fabric sits underneath it.

NVLink: fast, and only inside the box

NVLink is NVIDIA's direct GPU-to-GPU link, with NVSwitch extending it so every GPU in a server can talk to every other at full speed. On Hopper it reaches up to 900 GB/s per GPU, an order of magnitude beyond PCIe.

The limitation that matters: NVLink does not cross machines. It connects GPUs inside one server, or inside one rack-scale system such as GB200 NVL72. The moment your job spans two servers, NVLink stops being part of the answer.

Which cards have it. This is a real purchasing consideration and it is easy to get wrong:

  • SXM cards have full NVLink. H100 SXM, H200 SXM, A100 SXM, B200, B300.
  • PCIe variants have limited or no NVLink, and fall back to PCIe for most GPU-to-GPU traffic.
  • The L40S has no NVLink at all, and neither does the RTX 4090 or RTX 5090. NVIDIA removed NVLink from the GeForce line with Ada Lovelace.

If you are choosing between card variants for multi-GPU work, this gap usually matters more than the per-card specifications. Our L40S vs H100 vs A100 comparison covers which cards have NVLink and MIG alongside the rest of the specs.

InfiniBand: the specialist fabric between servers

InfiniBand was built for high-performance computing and carries two features that matter for distributed training.

RDMA. Remote Direct Memory Access lets one machine read and write another machine's memory without involving either CPU. That removes a copy and a context switch from every transfer, which is where much of the latency advantage comes from.

Lossless by design. InfiniBand uses credit-based flow control, so a sender does not transmit unless the receiver has buffer space. Packets are not dropped under congestion, so there is nothing to retransmit. Latency stays predictable when the fabric is busy, which is exactly when a synchronous training job is most sensitive to it.

Typical InfiniBand latency is in the low single-digit microseconds, with links commonly at 200 or 400 Gb/s.

Ethernet, and why RoCE changed the comparison

Standard Ethernet was not designed for this. It is lossy by default, and TCP handles congestion by dropping packets and retransmitting, which is fine for web traffic and poor for an all-reduce operation where every GPU waits for the slowest transfer.

RoCE closes most of the gap. RDMA over Converged Ethernet brings RDMA to Ethernet, and RoCEv2 is routable across subnets. Paired with a lossless Ethernet configuration using Priority Flow Control and Explicit Congestion Notification, it delivers much of what InfiniBand offers on more familiar hardware.

The catch is that lossless Ethernet must be configured correctly, and it is easy to get wrong. PFC misconfiguration can cause head-of-line blocking or, in the worst case, a deadlock across the fabric. InfiniBand gives you this behaviour by default; Ethernet gives it to you if your network team builds it properly.

We have a deeper treatment of this specific trade-off in RoCE vs InfiniBand for multi-node GPU training, which is worth reading if you are choosing between them for a build.

What each one costs

Honest answer first: we cannot give you reliable per-port pricing, because InfiniBand and Ethernet switch and adapter pricing is quoted through vendors and channel partners rather than published, and it moves. Anyone quoting you a precise figure in an article is guessing.

What is reliably true about the cost structure:

  • InfiniBand carries a hardware premium. Adapters, switches and cables are specialised and largely single-vendor, which is reflected in the price.
  • Ethernet has a broader supply base, more vendors, and hardware your team probably already knows.
  • The hidden cost of Ethernet is expertise. A correctly tuned lossless Ethernet fabric requires skills that are scarcer than the hardware saving. Budget for the engineering, not just the switches.
  • Renting removes the question entirely. If you use Runpod Clusters, the fabric is already built and priced into the GPU hour. There is no capital outlay and no network team required.

When you genuinely need a fast fabric

Not every distributed job needs InfiniBand-class networking. The deciding factor is how often your GPUs have to talk to each other.

You need it when:

  • You are running synchronous data-parallel training across nodes. Every step ends in an all-reduce where each GPU waits for all the others. This is the classic case, and the one where interconnect dominates.
  • You are sharding a model across machines. Tensor or pipeline parallelism moves activations between nodes constantly, not just gradients.
  • You are scaling past a handful of nodes. Communication overhead grows with node count, so the larger the cluster, the more the fabric decides your scaling efficiency.

You probably do not need it when:

  • Everything fits in one machine. Eight GPUs in a single node with NVLink is a lot of capacity, and no cluster fabric is involved. Most fine-tuning work never leaves this territory.
  • Your work is embarrassingly parallel. Hyperparameter sweeps, batch inference and independent experiments barely communicate. See distributed hyperparameter search for that pattern.
  • You can trade communication for computation. Gradient accumulation reduces synchronisation frequency at the cost of doing more work between syncs, which can make a slower fabric perfectly workable.

Before you blame the network

Interconnect gets blamed for a lot of problems it did not cause. Check these first, because they are more common and cheaper to fix.

  1. Is the GPU actually waiting on the network, or on data? A starved data pipeline looks like poor scaling. If utilisation swings between 100% and near zero, the problem is upstream. Our guide on eliminating GPU idle time from slow data pipelines covers how to tell.
  2. Are you using mixed precision? FP16, BF16 or FP8 reduce the volume of data crossing the network as well as speeding up compute. See mixed precision training.
  3. Is your collective library configured properly? NCCL usually picks sensible defaults, but a misdetected topology can route traffic over PCIe when NVLink was available.
  4. Do you actually need multiple nodes? A single node with more or larger GPUs is simpler and often cheaper than a cluster. Worth pricing before you scale out.

How this works on Runpod

Runpod Clusters scale to 64 GPUs with high-speed networking already in place, cited in our documentation at 1600 to 3200 Gbps between nodes. There is no fabric to procure, configure or tune, and no separate networking charge, because there are no ingress or egress fees on Runpod.

Within a node, SXM cards carry NVLink. Across nodes, the cluster fabric handles it. Clusters are self-serve with no commitment, so testing whether your job actually scales costs the hours it runs rather than a hardware purchase. See current pricing, or the Instant Clusters guide for the research workflow.

For the wider picture on scaling beyond one card, see the complete guide to multi-GPU training and GPU cluster management.

FAQ

Is NVLink faster than InfiniBand?

Yes, but they are not alternatives. NVLink on Hopper reaches up to 900 GB/s between GPUs in the same server, which is roughly 7,200 Gb/s and far beyond any cluster fabric. It simply does not work between machines. Use NVLink inside a node and InfiniBand or Ethernet between nodes. A well-built cluster uses both.

What is the difference between InfiniBand and Ethernet for AI training?

InfiniBand is lossless by design, using credit-based flow control so packets are not dropped under congestion, and it delivers RDMA natively at low single-digit microsecond latency. Ethernet is lossy by default, but RoCEv2 plus a correctly configured lossless fabric closes most of the gap. The practical difference is less about peak numbers and more about behaviour under load, and about whether you have the expertise to configure Ethernet properly.

Can I use Ethernet instead of InfiniBand for distributed training?

Often, yes. High-speed Ethernet with RoCEv2 is a legitimate choice, and many large clusters run on it. It becomes the wrong choice when your workload is heavily synchronous at large node counts, or when nobody on your team can build and maintain a lossless fabric. If you are renting rather than building, the question is usually moot because the provider has already made the decision.

Do I need InfiniBand if all my GPUs are in one server?

No. Inside a single machine, GPUs communicate over NVLink where available and PCIe otherwise. A cluster fabric is only relevant once you span multiple servers. If you are choosing hardware for a single node, prioritise SXM cards with NVLink over PCIe variants.

What is RoCE?

RDMA over Converged Ethernet. It brings the direct-memory-access model of InfiniBand to Ethernet hardware, with RoCEv2 adding routability across subnets. It needs a lossless Ethernet configuration to perform well, which is where most of the implementation difficulty lives.

How much does InfiniBand improve training performance?

It depends entirely on how communication-bound your job is, and any single percentage figure should be treated with suspicion. A synchronous data-parallel job across many nodes can be transformed by it. An embarrassingly parallel workload will see almost nothing. Measure your own job at two node counts and look at the scaling efficiency rather than trusting a benchmark run on someone else's model.

Get started

The reliable way to answer this for your own workload is to run it. Runpod Clusters are self-serve, scale to 64 GPUs, and bill by the second with no commitment. Launch a cluster or see current pricing.

Author profile: The Runpod Team

Purple glow background

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.

Star field background