News icon

Kimi K3 is now available on Runpod

L40S vs H100 vs A100: Choosing the Right GPU for AI

Most people arriving at this question are really asking one of two things: will my model fit, and am I paying for capability I do not need. Those are the questions this guide answers, for the three cards people compare most often and for the newer hardware that has since arrived above them.

Runpod rates below are pulled live from our pricing page. NVIDIA specifications were checked against NVIDIA's own product documentation on 31 August 2026.

The short answer

If you are doing thisStart hereRunpod hourly price
Inference on models up to roughly 13B, image and video generationL40S$1.09/hr
Fine-tuning with LoRA or QLoRAL40S or A100 PCIeL40S: $1.09/hr; A100 PCIe: $1.59/hr
Full fine-tuning and training, single nodeH100 SXM$3.49/hr
Long context, or models that will not fit in 80GBH200$4.59/hr
Frontier-scale trainingB200 or B300B200: $6.79/hr; B300: $7.89/hr
Development, experiments, small modelsRTX 4090 or RTX A5000RTX 4090 Community: $0.34/hr; RTX A5000 Secure: $0.27/hr

L40S vs H100: the comparison people actually make

These two get compared constantly, and the honest answer is that they are built for different jobs rather than being fast and slow versions of the same thing.

The L40S is an inference and graphics card. Ada Lovelace architecture, 48GB of GDDR6 with ECC, 864 GB/s of memory bandwidth, 91.6 TFLOPS FP32, and fourth-generation Tensor Cores with FP8 support at 733 TFLOPS dense. It also carries RT Cores, which matter if your workload touches rendering or video.

The H100 is a training card. Hopper architecture, 80GB of HBM3, and on the SXM variant roughly 3.35 TB/s of memory bandwidth. That is close to four times the L40S figure, and memory bandwidth is what most large-model training is actually limited by.

Two differences decide most real decisions, and neither is a headline number:

  • The L40S has no NVLink. Multi-GPU communication happens over PCIe Gen4. The H100 SXM has NVLink. If you are training across more than one GPU, this gap matters more than the per-card specs, because slow inter-GPU communication means eight GPUs can deliver the effective throughput of four.
  • The L40S has no MIG support. You cannot partition it into smaller isolated instances. The H100 and A100 both support Multi-Instance GPU.

When the L40S is the better buy: single-GPU inference, image and video generation, LoRA fine-tuning, and anything where 48GB is enough. At $1.09/hr against $3.49/hr for an H100 SXM, you are paying a fraction for a card that is not the bottleneck on those workloads.

When the H100 is worth the difference: full-parameter training, multi-GPU work, long context windows, and any job where you are already fighting the 48GB ceiling.

A100 vs H100: is the older card still worth it?

Often, yes. The A100 is Ampere rather than Hopper, and the single most important difference is that the A100 has no FP8 support. On modern transformer stacks tuned for FP8, that gap is significant. The H100's Transformer Engine is the reason for most of the performance difference people quote.

The A100 offers 40GB or 80GB of HBM2e, with the 80GB SXM variant around 2 TB/s of memory bandwidth against the H100 SXM's 3.35 TB/s.

The case for the A100 is cost per finished run, not speed. An A100 PCIe is $1.59/hr and an A100 SXM is $1.59/hr, against $2.89/hr and $3.49/hr for the H100 equivalents. If your stack does not use FP8, much of what you would pay extra for goes unused. Benchmark your own code before assuming the newer card wins on cost.

A100 vs L40S: 80GB of HBM against 48GB of GDDR

This comes down to memory type as much as memory size. The A100's HBM2e delivers roughly 2 TB/s; the L40S's GDDR6 delivers 864 GB/s. For memory-bound training, that is the whole comparison.

For inference, the picture inverts. The L40S has FP8 support that the A100 lacks, it costs $1.09/hr against $1.59/hr, and 48GB is enough for most models people actually serve. Unless you need the memory or the bandwidth, the A100 is the more expensive way to do the same job.

Full specification comparison

GPUArchitectureMemoryBandwidthFP8NVLinkMIGSecure Cloud
B300Blackwell Ultra288GB HBM3eHighest availableYes, plus FP4YesYes$7.89/hr
B200Blackwell180GB usable~8 TB/sYes, plus FP4YesYes$6.79/hr
H200 SXMHopper141GB HBM3e~4.8 TB/sYesYesYes$4.59/hr
H100 SXMHopper80GB HBM3~3.35 TB/sYesYesYes$3.49/hr
H100 PCIeHopper80GB HBM3~2 TB/sYesLimitedYes$2.89/hr
A100 SXMAmpere80GB HBM2e~2 TB/sNoYesYes$1.59/hr
A100 PCIeAmpere80GB HBM2e~1.9 TB/sNoLimitedYes$1.59/hr
RTX PRO 6000Blackwell96GBSee NVIDIA specsYesNoYes$2.09/hr
L40SAda Lovelace48GB GDDR6 ECC864 GB/sYes, 733 TFLOPS denseNoNo$1.09/hr
RTX 6000 AdaAda Lovelace48GBSee NVIDIA specsYesNoNo$0.84/hr
A40Ampere48GBSee NVIDIA specsNoYesNo$0.49/hr
RTX 5090Blackwell32GBSee NVIDIA specsYesNoNo$0.99/hr
RTX 4090Ada Lovelace24GBSee NVIDIA specsYesNoNo$0.74/hr
AMD MI300XCDNA 3192GB HBM3~5.3 TB/sROCm stackInfinity FabricYes$2.39/hr

L40S figures verified against NVIDIA product documentation, 31 August 2026. AMD MI300X rate verified 31 August 2026. All Runpod NVIDIA rates are live.

What replaced each of these cards

A common search is for the successor to a card you already know. The honest answers:

  • The A100's successor is the H100, and for memory-bound work the H200. The A100 remains in wide use because it is cheaper per hour and still capable, not because nothing replaced it.
  • The H100's successor is the H200 within Hopper, which is the same compute with 141GB and considerably more bandwidth, and then the B200 in Blackwell.
  • The L40S sits alongside the RTX PRO 6000 Blackwell, which offers 96GB against the L40S's 48GB and a newer architecture. It is available here at $2.09/hr, with MIG partitions at 48GB and 24GB if you want a slice rather than the whole card.

Blackwell shipped some time ago, so treat any guide still describing it as forthcoming as out of date. B200 and B300 are both self-serve here.

How to choose, by constraint

Start with VRAM, because it is a hard limit

Everything else is a matter of speed. VRAM is a matter of whether the job runs at all. For training, budget roughly 16 bytes per parameter for mixed-precision Adam before activation memory, which is why a 7B model needs far more than 7GB to train and why LoRA and QLoRA exist.

Then check whether you are memory-bound or compute-bound

Profile before upgrading. If GPU utilization swings between 100% and near zero, your data pipeline is starving the GPU and a faster card will not help. If utilization is consistently high and epochs are slow, more compute is the answer.

Then check the interconnect, if you are using more than one GPU

This is the most commonly skipped step and the most expensive to get wrong. SXM cards with NVLink communicate far faster than PCIe cards. Above a single machine, the network between nodes matters more than the card inside them. Runpod Clusters scale to 64 GPUs with 1600 to 3200 Gbps between nodes.

Then check the billing model

A cheaper card billed by the hour can cost more than a faster card billed by the second, if your work runs in bursts. Runpod bills per second across Pods, Serverless and Clusters, with no ingress or egress fees.

Running at scale

If you are choosing GPUs for an organization rather than a project, three things matter beyond the card:

  • Capacity without a queue. Inventory is visible in the console at deployment time across 31 global regions, rather than behind a quota request.
  • Compliance evidence. Runpod is SOC 2 Type II certified and HIPAA and GDPR compliant. Reports, Business Associate Agreements and Data Processing Agreements are available for security review.
  • Committed capacity. Savings Plans and Reserved Clusters exist for predictable workloads, alongside on-demand with no commitment.

FAQ

Which is better for LLM inference, L40S or H100?

For models that fit in 48GB, the L40S usually wins on cost per request. It supports FP8, which most modern inference stacks use, and costs $1.09/hr against $3.49/hr. The H100 pulls ahead when the model does not fit, when you need very long context, or when you are serving enough concurrent users that 80GB of KV cache becomes the constraint.

Is the A100 still worth using in 2026?

Yes, for the right workload. It has no FP8 support, so on modern transformer stacks it lags Hopper meaningfully. But at $1.59/hr it is often the lowest total cost for a training run that finishes, particularly if your code does not use FP8 anyway. Benchmark rather than assume.

Can the L40S train large models?

It can train, but 48GB and the absence of NVLink limit how far. LoRA and QLoRA fine-tuning up to roughly 30B works well. Full-parameter training of large models is where you want 80GB or more and NVLink between cards.

What is the difference between H100 PCIe and H100 SXM?

Same 80GB of HBM3, different form factor and interconnect. SXM has higher memory bandwidth, around 3.35 TB/s against roughly 2 TB/s, and full NVLink for multi-GPU communication. PCIe is cheaper at $2.89/hr against $3.49/hr and is fine for single-GPU work. If you are training across several GPUs, take SXM.

Do I need Blackwell?

Only if you are training at frontier scale or need the memory. B200 offers 180GB usable and FP4 support, and B300 goes further again. For the large majority of fine-tuning and inference work, an H100 or an L40S is the sensible choice and the money saved buys more GPU hours.

What about AMD?

The MI300X carries 192GB of HBM3, more memory per GPU than any NVIDIA card at a comparable rate, at $2.39/hr on Secure Cloud in configurations up to eight GPUs. The requirement is that your stack runs on ROCm rather than CUDA. If it does, the memory capacity is difficult to ignore.

Get started

The fastest way to settle a GPU choice is to run your own code on both. Runpod bills by the second with no minimum and no quota request, so a real benchmark costs very little. See current pricing, or deploy a Pod.

Related comparisons

Author profile: The Runpod Team

Purple glow background

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.

Star field background