News icon

Kimi K3 is now available on Runpod

Best GPU for ComfyUI: What the benchmarks actually show in 2026

The usual advice for ComfyUI is to buy the biggest GPU you can afford. Our own benchmark data says that is wrong, and expensively so.

On SDXL workflows, an RTX 5090 generated images faster than an H100 SXM while costing roughly a third as much per hour. On the same workflow, an RTX A5000 produced images for about a quarter of the H100’s cost per image, and a B300 – the fastest card we tested – cost 6.6 times more per image than that A5000 to be 4.4 times quicker.

This guide covers what actually decides ComfyUI performance, how much VRAM each model family really needs, which card fits which workflow, and what happens when several people share one pod. Every claim comes from measured results rather than specifications.

Every figure comes from a benchmark run on a rented Runpod Secure Cloud pod between July 3 and Sept. 10, 2026. The timer runs on the pod against the local ComfyUI server, so no network time is included. Comparisons are only made within a group of runs that shared an identical configuration, so any difference between two rows in the same table is the GPU rather than the setup.

A note on the prices in the tables below. Hourly rates and cost-per-image figures are Secure Cloud list prices as of Sept. 10, 2026, and every cost is derived from the rate beside it. They are a snapshot of that date, not a live feed, so the two columns always agree with each other. For today’s rates, see the pricing page.

What actually decides ComfyUI speed

Two main factors determine ComfyUI’s speed, in this order.

First, does the model fit in VRAM. If the model and its working memory do not fit, ComfyUI does not refuse. It keeps the weights in the server’s system memory and copies pieces onto the card for every step of every image. The workflow runs, but the card spends its time waiting on memory rather than computing, and the result depends heavily on the machine the GPU is plugged into.

Second, memory bandwidth and compute. Once the model fits, generation speed tracks how fast the card can move weights and do the arithmetic. This is where newer consumer architectures do surprisingly well, because diffusion workloads at batch size one do not benefit much from the features that make datacenter cards expensive.

What does not decide it: NVLink, MIG, or multi-GPU support. A single ComfyUI generation runs on one GPU. Those features matter for distributed training and mean nothing here.

How much VRAM you need for ComfyUI

This is the most useful table in the guide, and the most likely to save you money. Peak VRAM measured during real runs:

WorkflowPeak VRAM measuredWhat that means for card choice
SDXL, including Juggernaut XL v96.7 to 7.4 GB16 GB is comfortable. Even 12 GB works
FLUX.2 Klein 9B, bf1613.1 to 26.2 GBRuns on 16 GB at a usable speed
Qwen-Image 2512, fp820.7 to 20.9 GB24 GB
Qwen-Image 2512, bf1620.5 to 46.8 GBNeeds roughly 46 GB to fit. Below that it runs from system memory

Those ranges are wide, and the reason differs by model. On FLUX.2 Klein, peak memory tracks whatever card it is given – 13.1 GB on a 16 GB card, 26.1 GB on a 96 GB one – and the small cards still finish in a reasonable time.

On Qwen-Image bf16 the wide range means something different. That model needs around 46 GB to load. A card that cannot hold it still runs the workflow, but from system memory, so the peak VRAM figure reflects the card’s capacity rather than the model’s appetite, and the time reflects the host as much as the GPU.

So a peak VRAM number on its own does not tell you whether a model fits. What it does tell you is that a workflow will usually run on less than you expect – sometimes at a perfectly good speed, sometimes not.

The SDXL number is the one worth sitting with. A Juggernaut XL v9 workflow peaked at around 7 GB across every card from a 16 GB RTX 2000 Ada to a 288 GB B300. People routinely rent 48 GB and 80 GB cards for this. You are not buying capability at that point, only speed.

SDXL and Juggernaut XL: sixteen cards, measured

The widest comparison in the dataset, and the workflow most ComfyUI users are actually running. Sorted fastest first.

GPUVRAMSeconds per imageCost per imageSecure Cloud rate
B300288 GB2.26 s$0.00495$7.89/hr
RTX PRO 6000 Blackwell SE96 GB2.26 s$0.00131$2.09/hr
RTX 509032 GB3.04 s$0.00084$0.99/hr
H100 SXM80 GB3.43 s$0.00333$3.49/hr
L40S48 GB4.55 s$0.00138$1.09/hr
RTX PRO 6000 SE MIG 2g.48gb48 GB4.64 s$0.00141$1.09/hr
RTX 409024 GB4.79 s$0.00098$0.74/hr
RTX PRO 450032 GB6.22 s$0.00124$0.72/hr
RTX PRO 400024 GB7.35 s$0.00116$0.57/hr
A4048 GB7.41 s$0.00101$0.49/hr
A100 SXM80 GB7.65 s$0.00338$1.59/hr
RTX 309024 GB8.91 s$0.00124$0.50/hr
RTX A500024 GB9.95 s$0.00075$0.27/hr
RTX PRO 6000 SE MIG 1g.24gb24 GB10.13 s$0.00166$0.59/hr
RTX A400016 GB12.29 s$0.00085$0.25/hr
RTX 2000 Ada16 GB17.70 s$0.00118$0.24/hr

Juggernaut XL v9 in ComfyUI, single user, identical configuration across all sixteen cards. Median of repeat runs. Rates and costs are Secure Cloud list as of Sept. 10, 2026.

Three things stand out.

The RTX 5090 is third fastest and beats the H100 SXM, which costs more than three times as much per hour. On a batch-size-one diffusion workflow, the H100’s advantages in memory capacity and interconnect do not apply.

The two fastest cards tie at 2.26 seconds despite a threefold difference in hourly rate. A B300 has 288 GB of memory and an RTX PRO 6000 SE has 96 GB, and on a workflow that peaks at 7.4 GB neither is using what it charges for. The B300 costs $0.00495 per image; the RTX PRO 6000 SE costs $0.00131 for the same speed.

The RTX A5000 is the cheapest per image at roughly a quarter of the H100’s cost, despite taking more than three times as long. If you are generating in bulk and not watching each image appear, that trade is worth taking. If you are iterating interactively, it is not.

FLUX.2 Klein 9B: sixteen cards

GPUVRAMSeconds per imageCost per imagePeak VRAM used
H100 SXM80 GB1.47 s$0.0014226.1 GB
RTX PRO 6000 Blackwell SE96 GB1.51 s$0.0008826.1 GB
B300288 GB1.58 s$0.0034526.2 GB
RTX 509032 GB2.28 s$0.0006326.0 GB
RTX PRO 6000 SE MIG 2g.48gb48 GB3.18 s$0.00096not recorded
A100 SXM80 GB3.27 s$0.0014425.8 GB
L40S48 GB3.31 s$0.0010025.8 GB
RTX PRO 450032 GB3.40 s$0.0006825.7 GB
RTX 409024 GB3.79 s$0.0007821.1 GB
A4048 GB4.80 s$0.0006525.6 GB
RTX PRO 400024 GB5.04 s$0.0008021.0 GB
RTX A500024 GB6.46 s$0.0004921.1 GB
RTX 309024 GB6.81 s$0.0009521.1 GB
RTX PRO 6000 SE MIG 1g.24gb24 GB8.14 s$0.00133not recorded
RTX A400016 GB9.42 s$0.0006513.2 GB
RTX 2000 Ada16 GB13.19 s$0.0008813.1 GB

FLUX.2 Klein 9B bf16 in ComfyUI, single user, official-template four-step configuration. Cost per image uses Secure Cloud list prices as of Sept. 10, 2026.

Here the H100 does win on speed, by 40 milliseconds over an RTX PRO 6000 SE that costs 40% less per hour. But the RTX A5000 produces images for a third of the H100’s cost, and the RTX 5090 for less than half, in both cases at a fraction of the hourly rate.

The two 16 GB cards are the story in this table. FLUX.2 Klein is widely described as needing 24 GB. It ran on an RTX A4000 and an RTX 2000 Ada, peaking at 13.2 and 13.1 GB. They are the slowest cards here, but the gap is a matter of seconds rather than the collapse you see when a model genuinely does not fit, and at $0.00065 the RTX A4000 is close to the cheapest per image in the table.

The MIG partitions are worth understanding rather than dismissing: a 24 GB slice of an RTX PRO 6000 is slower than a whole 24 GB card, because you are sharing the underlying silicon. MIG exists to make a large card divisible, not to be fast.

Qwen-Image 2512: where VRAM starts to bite

GPUVRAMSeconds per imageCost per imagePeak VRAM used
B300288 GB9.35 s$0.0205046.8 GB
H100 SXM80 GB12.72 s$0.0123346.4 GB
RTX PRO 6000 Blackwell SE96 GB18.03 s$0.0104746.4 GB
A100 SXM80 GB29.12 s$0.0128646.1 GB
RTX 509032 GB31.81 s$0.0087528.7 GB
L40S48 GB38.91 s$0.0117841.4 GB
RTX PRO 6000 SE MIG 2g.48gb48 GB47.47 s$0.01437not recorded
RTX 6000 Ada48 GB52.76 s$0.0123144.6 GB
RTX PRO 450032 GB58.52 s$0.0117028.4 GB
A4048 GB59.66 s$0.0081241.8 GB
RTX 309024 GB86.82 s$0.0120620.5 GB
RTX 409024 GB108.90 s$0.0223820.7 GB
RTX A500024 GB132.70 s$0.0099521.2 GB

Qwen-Image 2512 bf16 in ComfyUI, single user. Cost per image uses Secure Cloud list prices as of Sept. 10, 2026. Cards under 46 GB run this model from system memory; their times depend on the host they landed on and should not be compared to each other. The fp8 variant is a separate configuration and is not directly comparable to these rows.

This is a heavier model, and the gap between cards widens accordingly. The A40 takes more than six times as long as the B300 but produces images for around 40% of the cost, which is the clearest illustration in the dataset that speed and cost efficiency are different questions.

The bottom of this table is not a ranking. Qwen-Image bf16 needs roughly 46 GB to load. Cards below that threshold do not refuse it, but they run it from the server’s system memory, and once a card is waiting on memory rather than computing, the number you get is set by the machine it is plugged into. That is why an RTX 3090 posted a faster time here than an RTX 4090. It is not a finding about either card.

Read those rows as one fact rather than five: below about 46 GB, this model is impractically slow. An RTX A5000 took over two minutes per image against the B300’s nine seconds. If you need Qwen-Image bf16 at a usable speed, you need a card it fits on.

The fp8 variant

If 24 GB is what you have, the fp8 build is the route that actually works.

GPUVRAMSeconds per imageCost per imagePeak VRAM used
RTX 409024 GB44.72 s$0.0091920.9 GB
RTX PRO 6000 SE MIG 2g.48gb48 GB51.82 s$0.01569not recorded
RTX PRO 400024 GB87.33 s$0.0138320.7 GB
RTX PRO 6000 SE MIG 1g.24gb24 GB118.67 s$0.01945not recorded

Qwen-Image 2512 fp8 in ComfyUI, single user. Cost per image uses Secure Cloud list prices as of Sept. 10, 2026. A separate configuration from the bf16 table, so these rows are comparable to each other and not to the ones above.

On the same RTX 4090, fp8 finished in 44.72 s against 108.90 s for bf16. The fp8 row is the 4090’s meaningful Qwen number, because the model fits in 24 GB and the card is doing the work rather than waiting on the host.

What happens when several people share one pod

Everything above measures one person working alone. A team sharing a pod is a different situation, and it is the one most studios actually run.

ComfyUI renders one image at a time. There is no batching across users and no overlapping of work, so a shared pod is simply a queue: your image waits for whatever is ahead of it, then renders. A card’s ceiling is one divided by its solo render time, and nothing you do to the queue raises it.

To measure what sharing feels like, we sent jobs at random intervals to an already-warm ComfyUI server at fixed fractions of each card’s capacity, then timed every request from submission to finished image, queue time included. Two levels are worth publishing: half capacity, and 90% of capacity.

Juggernaut XL under shared load

GPUAloneTypical wait, half busyTypical wait, 90% busySlowest 5%, 90% busyImages per hour
RTX PRO 6000 Blackwell SE2.26 s2.5 s7.4 s14.0 s1,487
RTX 50903.04 s3.5 s13.8 s25.3 s1,048
L40S4.55 s5.0 s18.4 s35.8 s694
RTX 40904.79 s5.7 s21.8 s41.5 s659
A407.41 s7.8 s27.1 s53.9 s455
A100 SXM7.65 s7.6 s23.1 s45.8 s445
RTX 30908.91 s10.6 s29.6 s64.1 s374
RTX A50009.95 s9.4 s38.7 s80.5 s336
RTX A400012.29 s12.1 s24.6 s72.7 s265
RTX 2000 Ada17.70 s17.6 s18.9 s43.0 s181

Juggernaut XL v9 in ComfyUI under offered load. Waits are measured from submission to finished image and include queue time. The RTX A4000 and RTX 2000 Ada completed fewer requests inside the measurement window than the other cards, so their 90% figures rest on a smaller sample and should be read as indicative.

FLUX.2 Klein under shared load

GPUAloneTypical wait, half busyTypical wait, 90% busySlowest 5%, 90% busyImages per hour
RTX PRO 6000 Blackwell SE1.51 s1.5 s5.2 s9.1 s2,206
RTX 50902.28 s2.5 s8.1 s15.9 s1,440
L40S3.31 s4.1 s12.2 s22.4 s866
RTX 40903.79 s4.0 s12.9 s26.5 s863
A404.80 s5.7 s17.8 s36.2 s673
RTX A50006.46 s7.2 s25.5 s50.5 s492
RTX 30906.81 s7.2 s25.0 s50.1 s493
RTX A40009.42 s9.5 s24.7 s57.8 s348
RTX 2000 Ada13.19 s13.0 s25.4 s74.0 s247

FLUX.2 Klein 9B bf16 in ComfyUI under offered load. Same measurement as the table above. The A100 SXM has no FLUX.2 Klein run under load; that job failed on a model download rather than on the benchmark.

What the numbers say

At half capacity, sharing is close to free. Across both models and every card tested, the typical wait with the server half busy lands within a quarter of the solo render time, and usually within a fifth. On several cards it is indistinguishable. A second person on the pod is not something the first person notices.

At 90% it changes completely. The typical wait runs three to four times the solo render, and one request in twenty waits six to eight times as long. On an RTX 4090 running Juggernaut XL, a 4.8-second render becomes a 21.8-second wait, with the slow tail at 41.5 seconds. That is the difference between a tool that feels responsive and one people complain about.

Queueing does not make images cheaper. This is the part worth knowing before you plan capacity. Because ComfyUI does not overlap work, running a pod at its ceiling produces no throughput bonus. Cost per image at full load came out higher than the single-user figure on all nineteen card-and-model combinations tested, typically by around 10%. Sharing a pod buys you utilization, not efficiency. What you spend is wait time.

The images-per-hour column is the number to plan against. It is the ceiling for one pod, and it follows directly from the solo render time. If your team needs 3,000 Juggernaut images a day, one RTX PRO 6000 SE covers it with room to spare; an RTX A5000 does not, however cheap each image is.

Two gaps worth stating plainly. The H100 SXM and B300 have no runs under load yet, so neither appears in these tables. And Qwen-Image renders slowly enough that too few requests completed inside the measurement window to compare cards fairly, so it is not published here.

Which GPU should you pick for ComfyUI?

The rates below are live, so they show today’s Secure Cloud price rather than the Sept. 10 snapshot used in the tables.

For SDXL and SDXL-derived checkpoints: an RTX 4090 at {{gpu:rtx-4090}}/hr, or {{gpu:rtx-4090:community}}/hr on Community Cloud, is the sensible default. The RTX 5090 at {{gpu:rtx-5090}}/hr is meaningfully faster if you are iterating and your time matters.

For bulk generation where you are not watching: an RTX A5000 at {{gpu:rtx-a5000}}/hr had the lowest cost per image of any card tested on SDXL, and on FLUX.2 Klein as well. Queue the work and let it run.

For FLUX.2 and similar models: 32 GB removes the pressure and the RTX 5090 is the value pick. If you already own 16 GB, try it before you upgrade – it runs.

For Qwen-Image bf16: you want a card the model fits on, which means roughly 46 GB or more. The RTX PRO 6000 Blackwell SE at {{gpu:rtx-pro-6000}}/hr is the cheapest card we tested that clears it. Below that threshold the model still runs, but from system memory, and the time you get stops being about the GPU.

If several people will share one pod: size against the images-per-hour ceiling rather than the solo render time, and keep the pod under roughly half its capacity if the work is interactive. Past that the wait grows much faster than the throughput does, and the cheapest card per image is rarely the right answer for a shared queue.

For learning and light experimentation: 16 GB is enough for SDXL and enough for FLUX.2 Klein. Start cheap and move up when speed rather than capacity becomes the problem.

When the biggest card is the right answer: rarely, on one image at a time. The B300 was fastest or joint fastest on every workflow here and most expensive per image on all three. It earns its rate on jobs that use its memory, which none of these workflows do.

Pods or Serverless for ComfyUI

Interactive work belongs on a Pod. You get the ComfyUI interface in a browser, persistent storage for models and outputs, and per-second billing.

API-driven or batch generation belongs on Serverless, which scales to zero between requests and costs nothing while idle. See deploying ComfyUI as a serverless API endpoint for that setup.

The shared-load numbers above are also an argument for Serverless over a shared Pod. A queue on one pod is a queue; separate workers are not. If several people are waiting on the same card for most of the day, you are paying in their time for a saving that the cost-per-image figures do not actually deliver.

Choosing for a team rather than a workstation

Everything above answers the question an individual asks: which card runs my workflow fastest for what it costs. Buying for a team changes the question, because you are no longer choosing for the workflow you run today. Three things shift.

Size for the heaviest workflow, not the typical one

The measurements above show SDXL peaking around 7 GB of VRAM. That is genuinely small, and on its own it argues for modest cards. But the Qwen-Image rows show what happens when someone submits something heavier than the card can hold: the job does not fail, it just becomes several times slower, and the slowdown depends on the host rather than on anything you chose.

For one person running one workflow, size for that workflow. For a shared queue running whatever an artist submits, size for the heaviest thing anyone will run, or jobs run slowly and unpredictably and nobody can tell you why. The practical version: pick your ceiling workflow first, find out how much memory it needs to load, then choose a card that holds it.

The card matters less than the hours you leave it running

This is the part that surprises people. Cost per image varies meaningfully across the cards above, but at team volumes the larger number is usually idle time, not hardware choice.

An artist works roughly eight hours. A pod left running bills for twenty-four. That is about two thirds of the spend buying nothing, on every seat, every day, whichever card you chose. Across a team the arithmetic gets uncomfortable quickly, and no amount of card optimization recovers it.

Two things address it directly. Runpod bills per second, so a pod stopped at the end of a session stops costing immediately rather than rounding up to the hour. And for anything that runs as a job rather than a session, a Serverless endpoint scales to zero between requests, so you pay for generation time instead of seat time.

That second option is worth understanding before you standardize a fleet, because it changes which card you should buy. If the work arrives as API calls rather than artists at keyboards, you are sizing for throughput and cold start rather than for a comfortable interactive session. We have covered that setup separately: see Deploy ComfyUI as a serverless API endpoint and the ComfyUI to API documentation.

Where the models and the outputs live

An individual keeps checkpoints on a local disk. A team cannot, and the reason is rarely convenience. Fine-tuned checkpoints and custom LoRAs trained on brand assets are usually the most valuable thing in the pipeline, and the question of where they sit tends to arrive from somebody outside the creative team.

A network volume keeps models in one place rather than re-downloading them onto every pod, which also removes several minutes from every cold start. On the compliance question: Runpod is SOC 2 Type II certified and HIPAA and GDPR compliant, with SOC 2 reports, BAAs, and DPAs available for security review through the Runpod Trust Center. Secure Cloud adds network isolation for workloads with stricter compliance needs. Coverage can vary by region and deployment model, so check requirements for your workload.

Related ComfyUI guides

Model-specific walkthroughs, each with the workflow already configured:

For Stable Diffusion outside ComfyUI, see the Stable Diffusion guide.

FAQ

Is an H100 worth it for ComfyUI?

For SDXL workflows, no. An RTX 5090 was faster in our benchmarks at roughly a third of the hourly rate, and an RTX PRO 6000 Blackwell SE was faster still. The H100 leads on FLUX.2 Klein by a small margin. Its advantages are memory capacity, interconnect and FP8 throughput, and a single-GPU diffusion workflow uses almost none of them.

How much VRAM do I need for ComfyUI?

It depends on the model, and there is a threshold rather than a gradient. SDXL peaked at around 7 GB on every card tested, and FLUX.2 Klein ran acceptably on 16 GB, so on those two, more VRAM buys speed rather than access. Qwen-Image bf16 is different: it needs roughly 46 GB to load, and cards below that run it from system memory several times slower. Find out what your heaviest model needs to load, then buy a card that holds it.

How many people can share one ComfyUI pod?

Fewer than you would guess, because ComfyUI renders one image at a time and gains nothing from a queue. Work out how many images an hour your team actually generates and compare it to the images-per-hour ceiling above. Keep the pod under about half that number and nobody notices they are sharing; push it to 90% and the typical wait triples while the slowest requests take six to eight times the solo render.

Does a 24 GB card work for FLUX?

Yes, comfortably, and so does a 16 GB card. FLUX.2 Klein peaked at 21.1 GB on the 24 GB cards and 13.2 GB on a 16 GB RTX A4000. Adding LoRAs or raising resolution will use more, so a 32 GB card gives you room to work, but 24 GB is not a constraint for the standard workflow.

Is the RTX 4090 still a good choice?

Yes. It handled every workflow we tested, sits near the top of the cost-per-image rankings on SDXL, and is the fastest card measured on Qwen-Image fp8. It is impractically slow on Qwen-Image bf16, but that is the model not fitting in 24 GB rather than a limitation of the card. The RTX 5090 is faster overall, and the 4090 remains a strong default.

Why is a MIG partition slower than a whole card of the same size?

A MIG slice is a partition of a larger GPU, so a 24 GB slice of an RTX PRO 6000 has a fraction of the full card’s compute, not all of it. MIG exists so one large card can serve several isolated workloads, which is useful for hosting many small jobs and unhelpful when you want one job to finish quickly.

Get started

Deploy a ComfyUI template from the Runpod Hub and it is running in a couple of minutes with no local setup. Billing is per second with no minimum, so testing two cards against your own workflow costs very little. See current pricing or deploy a Pod.

Purple glow background

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.

Star field background