News icon

Kimi K3 is now available on Runpod

GPU clusters for AI training: how to choose one in 2026

Comparing GPU clusters on per-GPU hourly rate alone hides the terms that decide the bill: a two-week minimum reservation, an orchestrator your job scripts do not run on, or an interconnect that makes a 16-GPU job slower than 8 GPUs would have been.

This guide works the decision in the order that actually matters: what the workload needs, what a cluster is made of, then who sells one on terms you can live with. Prices and terms below come from each provider's own pages, checked on October 5, 2026, with sources at the end.

TL;DR

  • Size the cluster from your model's parameter count and training method, not from a budget number. A full fine-tune needs roughly six times the VRAM of a LoRA on the same model.
  • Interconnect decides whether adding nodes helps. AWS and Google Cloud do not offer InfiniBand on their flagship training fabrics; they use EFA and RoCE respectively.
  • The commercial terms vary more than the hardware. Lambda requires a 16-GPU minimum and a two-week minimum reservation. CoreWeave has no public self-serve signup at all.

What a GPU cluster is, and when you need one

A GPU cluster is two or more GPU nodes wired together with a high-bandwidth, low-latency east-west network so a single job can span all of them. The network is the part that makes it a cluster rather than a pile of rented machines. On a standard 8-GPU node, that fabric runs at 400 Gbps per GPU, or 3,200 Gbps per node, on the H100-class InfiniBand clusters surveyed here.

Bandwidth is per GPU generation, not per provider. On Runpod Clusters, H100, H200 and B200 clusters run at 3,200 Gbps and A100 clusters at 1,600 Gbps. That halving matters more than it looks: on a synchronous job, the slowest link sets the step time for every GPU in the cluster, so an A100 cluster is not simply a cheaper H100 cluster with the same scaling behaviour.

Three workloads justify one.

Training from scratch or full fine-tuning a large model. This is the case where you genuinely have no choice. A 70B model in bf16 with Adam optimizer states does not fit in one node, so the job has to shard across nodes and the nodes have to talk constantly.

Fine-tuning where wall-clock time is the constraint. A LoRA on a 7B model fits on one GPU. Running twenty of them in parallel across a cluster is a throughput decision, not a memory one, and it is the cheapest kind of cluster job to schedule because the nodes barely need to talk.

Distributed inference for models too large for a single node. Rarer, and worth checking whether quantization or a smaller model gets you there on one node first, because a cluster raises your floor cost considerably.

If your job fits on a single 8-GPU node, rent a single node. Multi-node adds network failure modes, orchestration overhead, and a harder debugging story for no benefit.

Sizing: start from parameters, not budget

Memory per parameter depends almost entirely on training method. The arithmetic, in bytes per parameter:

MethodWeightsGradientsOptimizerTotal
Full fine-tune, bf16 + Adam221216
LoRA, bf16 base frozen2negligiblenegligible~2.5

Multiply by parameter count, then add 20 to 40% for activations and fragmentation. That gives a working floor:

ModelMethodVRAM floorFits on
7BLoRA~18 GB1 GPU, 24 GB class
7BFull fine-tune~112 GB2 GPUs, 80 GB class
70BLoRA~175 GB2 GPUs, 141 GB class
70BFull fine-tune~1,120 GB16 GPUs, 141 GB class
405BFull fine-tune~6,480 GB64+ GPUs, 141 GB class

The 70B full fine-tune row is the one that catches people. Eight 141 GB GPUs give you 1,128 GB, which clears 1,120 GB on paper and fails in practice once activations land. Sixteen is the honest answer.

These are floors for the memory math, not throughput predictions. How long a run takes depends on your batch size, sequence length, parallelism strategy, and how well your data pipeline keeps the GPUs fed. Benchmark your own step time on one node before you commit to a cluster size.

The six components to compare

ComponentWhat to askWhy it bites
GPU node hardwareVRAM per GPU, GPUs per node, generationVRAM per GPU sets the minimum node count; generation sets throughput
OrchestrationSlurm, Kubernetes, both, or neither, and is it managedA Slurm shop on a Kubernetes-only cluster rewrites its job scripts
NetworkingInfiniBand or Ethernet-based, Gbps per GPU, blocking or non-blockingDecides whether node 9 through 16 add throughput or just cost
StorageShared filesystem across nodes, local NVMe, throughputA checkpoint write that serializes across nodes stalls every GPU
Billing modelPer-second, per-minute, per-hour, or lump sumOn a 40-minute failed run, hourly billing costs 50% more than per-second
CommitmentOn-demand, minimum term, minimum sizeThe most common reason a cheap quote turns expensive

Billing granularity is the one most buyers skip and then regret. Distributed training fails often, early, and for boring reasons: a bad NCCL setting, a port conflict, a node that will not join. Ten failed 15-minute starts cost 10 hours on a per-hour meter and 2.5 hours on a per-second one.

What seven providers actually offer

All figures from each provider's own documentation, checked October 5, 2026.

Cluster size:

ProviderMin sizeMax on-demand
Runpod2 nodes, 16 GPUs64 GPUs self-serve
CoreWeaveNot published100k+ GPUs
Lambda16 GPUs512 or 2,000+, sources disagree
NebiusNot published32 GPUs self-serve
AWS1 instance64 instances per block
Google CloudNot publishedNot published
Azure1 VMNot published

Network and billing:

ProviderInterconnectBilling
Runpod3,200 Gbps on H100, H200 and B200, 1,600 on A100; InfiniBand or RoCE v2Per-second
CoreWeaveInfiniBand, 400 Gbps per GPU, 3.2 Tbps per nodeHourly rates listed
LambdaInfiniBand Quantum-2, 400 Gb/s per GPUHourly rates listed
NebiusInfiniBand Quantum-2; Quantum-X800 on B300Per-second
AWSEFA, not InfiniBandCapacity Blocks charged upfront
Google CloudRoCE or TCPXO, no InfiniBandNot published
AzureInfiniBand Quantum-2, 3.2 Tbps per VMPer-minute

Commitment and orchestration:

ProviderCommitmentOrchestration
RunpodNone on-demand; reserved 3mo+Managed Slurm or Ray; no Kubernetes
CoreWeaveOn-demand available; reserved term not publishedKubernetes; Slurm via SUNK
Lambda2 weeks minimumManaged Slurm or Kubernetes, pick one
NebiusNone stated for self-serveSoperator or Kubernetes, opt-in
AWSBlock reserved aheadNeither; ParallelCluster or EKS
Google CloudFlex-start has none; calendar mode reserves up to 90 daysNeither; GKE or Cluster Toolkit
AzureND-series excluded from capacity reservationsNeither; CycleCloud or AKS

Four things in that table decide most shortlists.

Two of the three hyperscalers have no InfiniBand on their training fabric. AWS uses its own Elastic Fabric Adapter and Google Cloud uses RoCE. The headline bandwidth can look comparable; the difference shows up in tail latency on collective operations, which is what synchronous training is bottlenecked on. Azure is the exception and ships native InfiniBand on ND H100 v5.

Lambda's floor is 16 GPUs and two weeks. That is the highest barrier to a first cluster in the survey, and it is a poor fit for the "does distributed training even help my job" question, which you answer in an afternoon.

CoreWeave does not sell to you without talking to you. Their own onboarding documentation says the sales team sends an invitation after approving your organization. Everything else about CoreWeave is built for scale, including a 100k+ GPU ceiling nobody else in this table advertises, so the gate is a deliberate segmentation choice rather than an oversight.

Nebius and Runpod are close peers on terms. Both bill GPU compute per second and both are self-serve. Runpod's self-serve ceiling is 64 GPUs and Nebius's is 32.

Modeling the cost

Published on-demand cluster rates from Runpod, per GPU per hour:

GPUVRAMRate
A100 SXM80 GB$1.79
H200 SXM141 GB$4.31

H100 SXM, B200 and L40S clusters, and every reserved cluster tier, are contact-sales with no published rate. If a published price is a hard requirement for your procurement process, that is worth knowing before you start a trial.

The arithmetic on the two published rates. A Runpod cluster starts at two nodes, so a single 8-GPU node is a Pod, not a Cluster, and is left out here:

ConfigurationPer hourPer 24 hours
16x A100 SXM, 2 nodes$28.64$687.36
32x A100 SXM, 4 nodes$57.28$1,374.72
16x H200 SXM, 2 nodes$68.96$1,655.04
32x H200 SXM, 4 nodes$137.92$3,310.08

Cross that against the sizing table. A 70B full fine-tune needs 16 GPUs at the 141 GB class, so $68.96 per hour is the floor for that job, or $1,655 for a day of training. The same job on a two-week Lambda minimum costs you the full two weeks whether the run takes three days or ten.

Availability is live and moves within the hour. In one Secure Cloud check on 2026-09-09, cluster-eligible GPU types with stock went from four to five and back inside two minutes, and the answer changes roughly sixfold depending on whether you ask for 1, 2, or 8 GPUs per node. Check the console at the moment you plan to launch rather than trusting any table, including this one.

Two sizing thresholds are worth knowing before you plan. Anyone can launch 2 nodes and up to 16 GPUs immediately. Going to 8 nodes and 64 GPUs needs a spend-limit increase rather than a sales contract, so it stays self-serve. Past 64 GPUs, up to 512, goes through sales, and reserved capacity scales beyond that.

When a Runpod cluster is the wrong choice

Four cases, stated plainly, because finding out later is expensive.

You run Kubernetes. Runpod Clusters ship managed Slurm and managed Ray and are not compatible with Kubernetes today. If your training stack assumes a Kubernetes control plane, CoreWeave is built around exactly that and this is not a close call. Ray users are fine; KubeRay users are not.

You need more than 64 GPUs without a sales conversation. Sixty-four is the self-serve ceiling, and it clears at 512 through sales and further on reserved capacity. CoreWeave and Lambda advertise higher ceilings if a large cluster today is the requirement.

You need a published price for an H100 or B200 cluster to get budget approved. Those tiers are contact-sales. Lambda publishes its 1-Click Cluster rates.

You need shared network storage in a specific region. Runpod Network Storage attaches to Clusters where the region offers it, so confirm storage availability in your target region before you plan a checkpointing strategy around it.

What this means

The hardware has largely converged. Runpod, CoreWeave, Lambda and Azure all quote 3,200 Gbps per 8-GPU node, and the GPUs are the same GPUs. What has not converged is the terms: whether you can start without a contract, whether a failed 20-minute run costs you 20 minutes or an hour, and whether you can find out the price without an email thread.

Decide the workload size first, from parameters and method. Then filter on orchestration, because that one is a rewrite if you get it wrong. Then compare terms. Rate per GPU-hour is the last filter, not the first, and it is the one that moves least between credible options.

Launch a 2-node Runpod Cluster and benchmark your own step time before committing to a size. For fine-tuning workloads that fit on one node, a single Pod is the cheaper start.

Sources

Purple glow background

Related articles

View All
No items found.

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.

Star field background