Comparing GPU clusters on per-GPU hourly rate alone hides the terms that decide the bill: a two-week minimum reservation, an orchestrator your job scripts do not run on, or an interconnect that makes a 16-GPU job slower than 8 GPUs would have been.
This guide works the decision in the order that actually matters: what the workload needs, what a cluster is made of, then who sells one on terms you can live with. Prices and terms below come from each provider's own pages, checked on October 5, 2026, with sources at the end.
TL;DR
- Size the cluster from your model's parameter count and training method, not from a budget number. A full fine-tune needs roughly six times the VRAM of a LoRA on the same model.
- Interconnect decides whether adding nodes helps. AWS and Google Cloud do not offer InfiniBand on their flagship training fabrics; they use EFA and RoCE respectively.
- The commercial terms vary more than the hardware. Lambda requires a 16-GPU minimum and a two-week minimum reservation. CoreWeave has no public self-serve signup at all.
What a GPU cluster is, and when you need one
A GPU cluster is two or more GPU nodes wired together with a high-bandwidth, low-latency east-west network so a single job can span all of them. The network is the part that makes it a cluster rather than a pile of rented machines. On a standard 8-GPU node, that fabric runs at 400 Gbps per GPU, or 3,200 Gbps per node, on the H100-class InfiniBand clusters surveyed here.
Bandwidth is per GPU generation, not per provider. On Runpod Clusters, H100, H200 and B200 clusters run at 3,200 Gbps and A100 clusters at 1,600 Gbps. That halving matters more than it looks: on a synchronous job, the slowest link sets the step time for every GPU in the cluster, so an A100 cluster is not simply a cheaper H100 cluster with the same scaling behaviour.
Three workloads justify one.
Training from scratch or full fine-tuning a large model. This is the case where you genuinely have no choice. A 70B model in bf16 with Adam optimizer states does not fit in one node, so the job has to shard across nodes and the nodes have to talk constantly.
Fine-tuning where wall-clock time is the constraint. A LoRA on a 7B model fits on one GPU. Running twenty of them in parallel across a cluster is a throughput decision, not a memory one, and it is the cheapest kind of cluster job to schedule because the nodes barely need to talk.
Distributed inference for models too large for a single node. Rarer, and worth checking whether quantization or a smaller model gets you there on one node first, because a cluster raises your floor cost considerably.
If your job fits on a single 8-GPU node, rent a single node. Multi-node adds network failure modes, orchestration overhead, and a harder debugging story for no benefit.
Sizing: start from parameters, not budget
Memory per parameter depends almost entirely on training method. The arithmetic, in bytes per parameter:
| Method | Weights | Gradients | Optimizer | Total |
|---|---|---|---|---|
| Full fine-tune, bf16 + Adam | 2 | 2 | 12 | 16 |
| LoRA, bf16 base frozen | 2 | negligible | negligible | ~2.5 |
Multiply by parameter count, then add 20 to 40% for activations and fragmentation. That gives a working floor:
| Model | Method | VRAM floor | Fits on |
|---|---|---|---|
| 7B | LoRA | ~18 GB | 1 GPU, 24 GB class |
| 7B | Full fine-tune | ~112 GB | 2 GPUs, 80 GB class |
| 70B | LoRA | ~175 GB | 2 GPUs, 141 GB class |
| 70B | Full fine-tune | ~1,120 GB | 16 GPUs, 141 GB class |
| 405B | Full fine-tune | ~6,480 GB | 64+ GPUs, 141 GB class |
The 70B full fine-tune row is the one that catches people. Eight 141 GB GPUs give you 1,128 GB, which clears 1,120 GB on paper and fails in practice once activations land. Sixteen is the honest answer.
These are floors for the memory math, not throughput predictions. How long a run takes depends on your batch size, sequence length, parallelism strategy, and how well your data pipeline keeps the GPUs fed. Benchmark your own step time on one node before you commit to a cluster size.
The six components to compare
| Component | What to ask | Why it bites |
|---|---|---|
| GPU node hardware | VRAM per GPU, GPUs per node, generation | VRAM per GPU sets the minimum node count; generation sets throughput |
| Orchestration | Slurm, Kubernetes, both, or neither, and is it managed | A Slurm shop on a Kubernetes-only cluster rewrites its job scripts |
| Networking | InfiniBand or Ethernet-based, Gbps per GPU, blocking or non-blocking | Decides whether node 9 through 16 add throughput or just cost |
| Storage | Shared filesystem across nodes, local NVMe, throughput | A checkpoint write that serializes across nodes stalls every GPU |
| Billing model | Per-second, per-minute, per-hour, or lump sum | On a 40-minute failed run, hourly billing costs 50% more than per-second |
| Commitment | On-demand, minimum term, minimum size | The most common reason a cheap quote turns expensive |
Billing granularity is the one most buyers skip and then regret. Distributed training fails often, early, and for boring reasons: a bad NCCL setting, a port conflict, a node that will not join. Ten failed 15-minute starts cost 10 hours on a per-hour meter and 2.5 hours on a per-second one.
What seven providers actually offer
All figures from each provider's own documentation, checked October 5, 2026.
Cluster size:
| Provider | Min size | Max on-demand |
|---|---|---|
| Runpod | 2 nodes, 16 GPUs | 64 GPUs self-serve |
| CoreWeave | Not published | 100k+ GPUs |
| Lambda | 16 GPUs | 512 or 2,000+, sources disagree |
| Nebius | Not published | 32 GPUs self-serve |
| AWS | 1 instance | 64 instances per block |
| Google Cloud | Not published | Not published |
| Azure | 1 VM | Not published |
Network and billing:
| Provider | Interconnect | Billing |
|---|---|---|
| Runpod | 3,200 Gbps on H100, H200 and B200, 1,600 on A100; InfiniBand or RoCE v2 | Per-second |
| CoreWeave | InfiniBand, 400 Gbps per GPU, 3.2 Tbps per node | Hourly rates listed |
| Lambda | InfiniBand Quantum-2, 400 Gb/s per GPU | Hourly rates listed |
| Nebius | InfiniBand Quantum-2; Quantum-X800 on B300 | Per-second |
| AWS | EFA, not InfiniBand | Capacity Blocks charged upfront |
| Google Cloud | RoCE or TCPXO, no InfiniBand | Not published |
| Azure | InfiniBand Quantum-2, 3.2 Tbps per VM | Per-minute |
Commitment and orchestration:
| Provider | Commitment | Orchestration |
|---|---|---|
| Runpod | None on-demand; reserved 3mo+ | Managed Slurm or Ray; no Kubernetes |
| CoreWeave | On-demand available; reserved term not published | Kubernetes; Slurm via SUNK |
| Lambda | 2 weeks minimum | Managed Slurm or Kubernetes, pick one |
| Nebius | None stated for self-serve | Soperator or Kubernetes, opt-in |
| AWS | Block reserved ahead | Neither; ParallelCluster or EKS |
| Google Cloud | Flex-start has none; calendar mode reserves up to 90 days | Neither; GKE or Cluster Toolkit |
| Azure | ND-series excluded from capacity reservations | Neither; CycleCloud or AKS |
Four things in that table decide most shortlists.
Two of the three hyperscalers have no InfiniBand on their training fabric. AWS uses its own Elastic Fabric Adapter and Google Cloud uses RoCE. The headline bandwidth can look comparable; the difference shows up in tail latency on collective operations, which is what synchronous training is bottlenecked on. Azure is the exception and ships native InfiniBand on ND H100 v5.
Lambda's floor is 16 GPUs and two weeks. That is the highest barrier to a first cluster in the survey, and it is a poor fit for the "does distributed training even help my job" question, which you answer in an afternoon.
CoreWeave does not sell to you without talking to you. Their own onboarding documentation says the sales team sends an invitation after approving your organization. Everything else about CoreWeave is built for scale, including a 100k+ GPU ceiling nobody else in this table advertises, so the gate is a deliberate segmentation choice rather than an oversight.
Nebius and Runpod are close peers on terms. Both bill GPU compute per second and both are self-serve. Runpod's self-serve ceiling is 64 GPUs and Nebius's is 32.
Modeling the cost
Published on-demand cluster rates from Runpod, per GPU per hour:
| GPU | VRAM | Rate |
|---|---|---|
| A100 SXM | 80 GB | $1.79 |
| H200 SXM | 141 GB | $4.31 |
H100 SXM, B200 and L40S clusters, and every reserved cluster tier, are contact-sales with no published rate. If a published price is a hard requirement for your procurement process, that is worth knowing before you start a trial.
The arithmetic on the two published rates. A Runpod cluster starts at two nodes, so a single 8-GPU node is a Pod, not a Cluster, and is left out here:
| Configuration | Per hour | Per 24 hours |
|---|---|---|
| 16x A100 SXM, 2 nodes | $28.64 | $687.36 |
| 32x A100 SXM, 4 nodes | $57.28 | $1,374.72 |
| 16x H200 SXM, 2 nodes | $68.96 | $1,655.04 |
| 32x H200 SXM, 4 nodes | $137.92 | $3,310.08 |
Cross that against the sizing table. A 70B full fine-tune needs 16 GPUs at the 141 GB class, so $68.96 per hour is the floor for that job, or $1,655 for a day of training. The same job on a two-week Lambda minimum costs you the full two weeks whether the run takes three days or ten.
Availability is live and moves within the hour. In one Secure Cloud check on 2026-09-09, cluster-eligible GPU types with stock went from four to five and back inside two minutes, and the answer changes roughly sixfold depending on whether you ask for 1, 2, or 8 GPUs per node. Check the console at the moment you plan to launch rather than trusting any table, including this one.
Two sizing thresholds are worth knowing before you plan. Anyone can launch 2 nodes and up to 16 GPUs immediately. Going to 8 nodes and 64 GPUs needs a spend-limit increase rather than a sales contract, so it stays self-serve. Past 64 GPUs, up to 512, goes through sales, and reserved capacity scales beyond that.
When a Runpod cluster is the wrong choice
Four cases, stated plainly, because finding out later is expensive.
You run Kubernetes. Runpod Clusters ship managed Slurm and managed Ray and are not compatible with Kubernetes today. If your training stack assumes a Kubernetes control plane, CoreWeave is built around exactly that and this is not a close call. Ray users are fine; KubeRay users are not.
You need more than 64 GPUs without a sales conversation. Sixty-four is the self-serve ceiling, and it clears at 512 through sales and further on reserved capacity. CoreWeave and Lambda advertise higher ceilings if a large cluster today is the requirement.
You need a published price for an H100 or B200 cluster to get budget approved. Those tiers are contact-sales. Lambda publishes its 1-Click Cluster rates.
You need shared network storage in a specific region. Runpod Network Storage attaches to Clusters where the region offers it, so confirm storage availability in your target region before you plan a checkpointing strategy around it.
What this means
The hardware has largely converged. Runpod, CoreWeave, Lambda and Azure all quote 3,200 Gbps per 8-GPU node, and the GPUs are the same GPUs. What has not converged is the terms: whether you can start without a contract, whether a failed 20-minute run costs you 20 minutes or an hour, and whether you can find out the price without an email thread.
Decide the workload size first, from parameters and method. Then filter on orchestration, because that one is a rewrite if you get it wrong. Then compare terms. Rate per GPU-hour is the last filter, not the first, and it is the one that moves least between credible options.
Launch a 2-node Runpod Cluster and benchmark your own step time before committing to a size. For fine-tuning workloads that fit on one node, a single Pod is the cheaper start.
Sources
- Runpod Clusters, Runpod pricing and Runpod Clusters documentation
- Lambda 1-Click Clusters and Lambda documentation
- CoreWeave networking and CoreWeave SUNK
- Nebius self-service and Nebius compute pricing
- AWS EC2 Capacity Blocks and AWS EFA
- Google Cloud GPU network bandwidth and provisioning models
- Azure ND H100 v5 and Azure capacity reservations
