Whisper barely touches VRAM. Across every configuration measured, peak usage runs between 1.1 GB and 4.4 GB, which means almost any rentable GPU can hold it. So the question is not what fits. It is what a transcribed hour costs, and on that the spread is enormous: $0.0071 to $0.5067 per hour of audio across 23 GPUs. Speed over the same range varies by about two.
Why cost per hour of audio is the only number that matters
Transcription is a batch job. Nobody watches it run. That makes it unusual among GPU workloads, because the metric everyone reaches for first, raw speed, decides almost nothing.
On Whisper large-v3-turbo with int8, the fastest card measured is an L40S at 58.3 times realtime. The slowest is an H200 at 27.0 times realtime. That is a 2.2x spread. Over the same 23 cards, cost per hour of audio ranges from $0.0071 to $0.1697, a 24x spread.
Put plainly: an hour of audio costs 24 times more on the wrong card and arrives roughly twice as fast. Unless you are transcribing live, that trade is not worth making.
The gap widens on the full large-v3 model, where a B300 transcribes at 15.6 times realtime for $0.5067 an hour of audio, while an RTX A5000 manages 19.4 times realtime for $0.0139. The cheaper card is both faster and 36 times less expensive.
Whisper large-v3-turbo with int8: speed and cost by GPU
The best-performing configuration measured, on 23 GPUs, sorted by cost per hour of audio.
| GPU | Card VRAM | Peak VRAM used | Realtime factor | $ per audio hour |
|---|---|---|---|---|
| RTX A5000 | 24 GB | 1.3 GB | 38.1x | $0.0071 |
| RTX 2000 Ada | 16 GB | 1.1 GB | 30.4x | $0.0079 |
| RTX A4000 | 16 GB | 1.2 GB | 30.5x | $0.0082 |
| RTX 3090 | 24 GB | 1.3 GB | 43.8x | $0.0114 |
| RTX A6000 | 48 GB | 1.1 GB | 42.8x | $0.0124 |
| A40 | 48 GB | 1.4 GB | 35.4x | $0.0139 |
| Pro 6000 MIG 24GB | 24 GB | not measured | 41.4x | $0.0143 |
| L4 | 24 GB | 1.1 GB | 31.5x | $0.0156 |
| L40 | 48 GB | 1.3 GB | 51.7x | $0.0159 |
| RTX 6000 Ada | 48 GB | 1.3 GB | 50.1x | $0.0168 |
| L40S | 48 GB | 1.5 GB | 58.3x | $0.0187 |
| RTX 5090 | 32 GB | 1.6 GB | 53.0x | $0.0187 |
| RTX 4090 | 24 GB | 1.5 GB | 35.2x | $0.0210 |
| Pro 6000 MIG 48GB | 48 GB | not measured | 48.0x | $0.0227 |
| A100 PCIe | 80 GB | 1.5 GB | 45.1x | $0.0353 |
| RTX Pro 6000 | 96 GB | 1.6 GB | 50.8x | $0.0411 |
| A100 SXM | 80 GB | 1.5 GB | 34.0x | $0.0467 |
| H100 NVL | 94 GB | 1.6 GB | 57.8x | $0.0552 |
| H100 SXM | 80 GB | 1.7 GB | 52.8x | $0.0661 |
| H100 PCIe | 80 GB | 1.6 GB | 34.3x | $0.0842 |
| B200 | 180 GB | 1.8 GB | 49.5x | $0.1371 |
| B300 | 288 GB | 1.8 GB | 47.9x | $0.1647 |
| H200 | 141 GB | 1.7 GB | 27.0x | $0.1697 |
Cost per hour of audio is calculated from Runpod Secure Cloud on-demand rates as of September 15, 2026. For current rates, see the pricing page.
Whisper large-v3 with int8: speed and cost by GPU
The full model on the same 23 GPUs, for cases where turbo’s accuracy is not enough.
| GPU | Peak VRAM used | Realtime factor | $ per audio hour |
|---|---|---|---|
| RTX A5000 | 2.1 GB | 19.4x | $0.0139 |
| RTX A4000 | 1.9 GB | 16.6x | $0.0151 |
| RTX 2000 Ada | 2.0 GB | 14.1x | $0.0170 |
| RTX A6000 | 1.8 GB | 20.8x | $0.0254 |
| RTX 3090 | 2.1 GB | 18.9x | $0.0265 |
| A40 | 2.1 GB | 16.6x | $0.0294 |
| L4 | 1.7 GB | 15.5x | $0.0316 |
| Pro 6000 MIG 24GB | not measured | 18.5x | $0.0319 |
| RTX 6000 Ada | 2.0 GB | 23.3x | $0.0361 |
| RTX 4090 | 1.9 GB | 19.7x | $0.0375 |
| L40 | 2.0 GB | 21.1x | $0.0388 |
| RTX 5090 | 2.4 GB | 24.8x | $0.0399 |
| L40S | 2.0 GB | 23.1x | $0.0472 |
| Pro 6000 MIG 48GB | not measured | 22.8x | $0.0479 |
| RTX Pro 6000 | 2.4 GB | 25.6x | $0.0818 |
| A100 PCIe | 2.1 GB | 18.1x | $0.0878 |
| A100 SXM | 2.3 GB | 16.6x | $0.0958 |
| H100 NVL | 2.4 GB | 24.4x | $0.1307 |
| H100 SXM | 2.4 GB | 24.9x | $0.1399 |
| H100 PCIe | 2.4 GB | 14.2x | $0.2032 |
| H200 | 2.4 GB | 22.4x | $0.2046 |
| B200 | 2.5 GB | 16.4x | $0.4145 |
| B300 | 2.5 GB | 15.6x | $0.5067 |
Which GPU should you pick for Whisper?
Cheapest GPU per hour of audio
The RTX A5000 at $0.0071 an hour of audio on turbo with int8, transcribing at 38.1 times realtime. An hour of recording takes about 95 seconds. For a thousand hours of archive, that is roughly $7 of compute.
The RTX 2000 Ada at $0.0079 and the RTX A4000 at $0.0082 sit within a fraction of a cent, both on 16 GB cards.
Fastest GPU for Whisper transcription
The L40S at 58.3 times realtime, with the H100 NVL at 57.8 and the RTX 5090 at 53.0. The L40S costs $0.0187 an hour of audio; the H100 NVL costs $0.0552 for a result 1% slower. On speed alone the L40S wins; on speed per dollar it is not close.
Best value: RTX 3090, RTX A6000 and L40
If you want speed without the cost floor, three cards land in the middle. The RTX 3090 at 43.8 times realtime for $0.0114, the RTX A6000 at 42.8 times for $0.0124, and the L40 at 51.7 times for $0.0159. All three are faster than the cheapest cards and still cost between a third and a tenth of what the H100 and Blackwell cards do per hour of audio.
Why datacenter GPUs lose on transcription
The H200, B200 and B300 are the three most expensive cards per hour of audio in this test, and none of them is the fastest. The H200 is the slowest of all 23 cards on turbo with int8, at 27.0 times realtime, while costing $0.1697 an hour of audio, 24 times an RTX A5000.
The reason is in the VRAM column. Whisper peaks at 1.1 to 2.5 GB. A B300 offers 288 GB. Nothing about the model uses what those cards are built to provide, so the hourly rate buys capability the workload cannot spend.
This is the opposite of LLM serving, where large VRAM and memory bandwidth translate directly into throughput. Transcription is one of the few GPU workloads where the cheap end of the catalog is the correct answer.
Does int8 quantization make Whisper faster?
Not reliably, and the direction depends on the model.
On large-v3, int8 is slower on 8 of the 10 cards measured in both. An H100 SXM drops from 27.6 to 24.9 times realtime, an RTX 4090 from 25.1 to 19.7, a B300 from 24.0 to 15.6. Only the L4 and the Pro 6000 MIG 24GB improve.
On turbo, int8 helps on 7 of 11. An A40 goes from 22.2 to 35.4 times realtime and an RTX A6000 from 28.0 to 42.8. But an RTX 4090 goes the other way, 49.3 down to 35.2.
What int8 does do consistently is roughly halve peak VRAM. On cards where the model already fits comfortably, that buys nothing. Test both on your chosen card rather than assuming quantization is a speedup.
How much VRAM does Whisper need?
Less than almost any other production model.
| Configuration | Peak VRAM measured |
|---|---|
| large-v3-turbo, int8 | 1.1–1.8 GB |
| large-v3-turbo, fp16 | 2.0–2.7 GB |
| large-v3, int8 | 1.7–2.5 GB |
| large-v3, fp16 | 3.3–4.4 GB |
These figures are lower than most published guidance. It is common to see Whisper large-v3 quoted at 10 GB of VRAM or more, and quantized builds at 8 GB. Measured across ten GPUs, large-v3 at fp16 peaks between 3.3 and 4.4 GB, and with int8 between 1.7 and 2.5 GB. If you have been sizing hardware from the higher figures, you have been provisioning two to three times more card than the model actually uses, and paying for it every hour.
Every card in this test has at least 16 GB. None came close to its limit. If you are choosing hardware for Whisper alone, VRAM should not enter the decision.
When a dedicated GPU is the wrong answer
If you transcribe occasionally, a few hours a week or in unpredictable bursts, a hosted transcription API will cost less than any GPU you rent by the hour, because you pay nothing between jobs. Serverless with scale-to-zero is the middle option.
Renting starts to win at volume, when the data cannot leave your control, or when you need a fine-tuned or non-English model that hosted providers do not serve.
What these benchmarks do not cover
- Accuracy is not measured. These are speed and cost figures only. Turbo trades some accuracy for speed, and int8 trades more. Whether that trade is acceptable depends on your audio and your tolerance for errors.
- Single-stream only. No concurrency or batching. Cost per hour of audio would fall on every card with parallel jobs, and probably fall furthest on the large cards, which have the most idle capacity here.
- English is assumed. Multilingual and translation modes are not separated out.
- faster-whisper only. WhisperX, whisper.cpp and the OpenAI API will differ.
- Peak VRAM is missing on the two Pro 6000 MIG partitions. Those cells say so rather than carrying an estimate.
- Cost is a snapshot from Secure Cloud rates on September 15, 2026.
Common questions about running Whisper on a GPU
What is the best GPU for Whisper?
For cost, an RTX A5000 transcribes an hour of audio for $0.0071 at 38.1 times realtime. For speed, an L40S reaches 58.3 times realtime at $0.0187. Both run Whisper large-v3-turbo with int8 quantization.
How much VRAM does Whisper need?
Between 1.1 GB and 4.4 GB depending on model and precision, which is less than commonly published guidance suggests. Whisper large-v3-turbo with int8 peaks at 1.1 to 1.8 GB; the full large-v3 at fp16 peaks at 3.3 to 4.4 GB. Figures of 8 to 10 GB are widely quoted and are roughly two to three times higher than these measurements. Any 16 GB card has ample headroom.
How fast is Whisper on a GPU?
Between 14 and 58 times realtime across the cards measured. An hour of audio takes roughly one to four minutes. Turbo is about twice as fast as large-v3 on the same hardware.
How much does it cost to transcribe an hour of audio?
Between $0.0071 and $0.5067 depending entirely on the card. A thousand hours costs about $7 on an RTX A5000 and about $507 on a B300, for a result that is slower.
Is Whisper large-v3-turbo worth it over large-v3?
On these measurements, yes for most uses. Turbo runs roughly twice as fast at roughly half the VRAM, and costs less per hour of audio on every card. The tradeoff is accuracy, which these benchmarks do not measure.
Should I use int8 quantization for Whisper?
Test it, do not assume. It reliably halves VRAM but does not reliably improve speed: on large-v3 it is slower on 8 of 10 cards measured, and on turbo it helps on 7 of 11. If the model already fits your card, the VRAM saving buys nothing.
Compare these figures interactively, alongside LLM, image and video workloads, on the GPU Compare tool.
Author profile: The Runpod Team
