
B300
Next-generation data center GPU based on NVIDIA Blackwell Ultra architecture with 288 GB HBM3e memory, delivering 1.5x FP4 performance and 2x attention performance over the B200 for AI reasoning and real-time inference workloads.
GPU Benchmarks
Compare performance across LLMs and image models to find the best GPU for your workload.

We have not run the Runpod benchmark suite on the B300 yet, so the charts below show the A40 and every other measured GPU without a B300 bar. The table gives our estimate instead: the measured B200 result for each workload, scaled by the one specification that governs it. The B300 (Blackwell Ultra) matches the B200 on memory bandwidth (8 TB/s), dense FP16 tensor throughput (2.25 PFLOPS) and dense FP8 (4.5 PFLOPS), so for BF16 and FP8 inference on 7B to 8B models the defensible estimate is parity with the B200. Where the B300 pulls ahead is 288 GB of HBM3e against 180 GB (1.6x, so larger models and longer contexts fit on one GPU) and 1.5x dense FP4 (15 vs 10 PFLOPS), and neither shows up in these BF16/FP8 tests on small models.
| Workload | A40 measured | B200 measured | B300 estimated | What sets the estimate |
|---|---|---|---|---|
| Llama 3.1 8B, 1,024 in / 1,024 out, batch 1 (tok/s) | 34 | 196 | 196 | Decode is memory-bandwidth-bound at this batch size: 8 TB/s on both parts |
| Llama 3.1 8B, 1,024 in / 1,024 out, batch 32 (tok/s) | 623 | 2,621 | 2,621 | Compute-bound at this batch size: dense FP8/FP16 tensor throughput is unchanged (4.5/2.25 PFLOPS) |
| Llama 3.1 8B, 1,024 in / 1,024 out, batch 128 (tok/s) | 919 | 3,928 | 3,928 | Compute-bound at this batch size: dense FP8/FP16 tensor throughput is unchanged (4.5/2.25 PFLOPS) |
| Qwen2.5 7B, 512 in / 512 out, batch 256 (tok/s) | 1,897 | 13,764 | 13,764 | Compute-bound at this batch size: dense FP8/FP16 tensor throughput is unchanged (4.5/2.25 PFLOPS) |
| Flux.1-dev, 512 px, 10 steps (s/image) | 3.9 | 1.1 | 1.1 | Diffusion is compute-bound in BF16: dense FP16 tensor throughput is unchanged (2.25 PFLOPS) |
| Flux.1-dev, 1,024 px, 50 steps (s/image) | 54.3 | 10.7 | 10.7 | Diffusion is compute-bound in BF16: dense FP16 tensor throughput is unchanged (2.25 PFLOPS) |
| SDXL 1.0, 1,024 px, 30 steps (s/image) | 6.7 | 4.4 | 4.4 | Diffusion is compute-bound in BF16: dense FP16 tensor throughput is unchanged (2.25 PFLOPS) |
How to read the estimates: expect measured B300 results within roughly 0.95x to 1.10x of the B200 figures. The low end is run-to-run noise; the high end is what the 1,400 W power envelope (against 1,000 W) and Blackwell Ultra's faster attention softmax could add on long-context runs. Values are output tokens per second for LLM rows (higher is better) and seconds per image for diffusion rows (lower is better). Sources: Runpod benchmark service measurements for the B200 and the comparison GPU; NVIDIA Blackwell and Blackwell Ultra specifications. Measured B300 results replace this table as soon as the runs land.
Benchmarks were run using vLLM in May 2025 with Runpod GPUs

Next-generation data center GPU based on NVIDIA Blackwell Ultra architecture with 288 GB HBM3e memory, delivering 1.5x FP4 performance and 2x attention performance over the B200 for AI reasoning and real-time inference workloads.

Data center GPU based on Ampere architecture with 48GB GDDR6 memory and 10,752 CUDA cores for AI workloads, professional visualization, and virtual workstation applications.

High-efficiency LLM processing at 90.98 tok/s.
Benchmarks were run using Hugging Face Diffusers in May 2025 on Runpod GPUs.

Unmatched image gen speed with 49.9 images per minute.

AI image processing at 40.3 images per minute.

Pro-grade performance with 36 images per minute.
Case Studies