News icon

Kimi K3 is now available on Runpod

RTX 4090
rtx-4090
RTX 3090
rtx-3090
RTX 4090
rtx-4090
RTX 2000 Ada
rtx-2000-ada
RTX 4090
rtx-4090
L40S
l40s
RTX 4090
rtx-4090
L40
l40
RTX 4090
rtx-4090
L4
l4
RTX 4090
rtx-4090
H100 SXM
h100-sxm
RTX 4090
rtx-4090
H100 PCIe
h100-pcie
RTX 4090
rtx-4090
H100 NVL
h100-nvl
RTX 4090
rtx-4090
A40
a40
RTX 4090
rtx-4090
A100 SXM
a100-sxm
RTX 4090
rtx-4090
A100 PCIe
a100-pcie
RTX 3090
rtx-3090
RTX A6000
rtx-a6000
RTX 3090
rtx-3090
RTX A5000
rtx-a5000
RTX 3090
rtx-3090
RTX A4000
rtx-a4000
RTX 3090
rtx-3090
RTX 6000 Ada
rtx-6000-ada
RTX 3090
rtx-3090
RTX 4090
rtx-4090
RTX 3090
rtx-3090
RTX 2000 Ada
rtx-2000-ada
RTX 3090
rtx-3090
L40S
l40s
RTX 3090
rtx-3090
L40
l40
RTX 3090
rtx-3090
L4
l4
RTX 3090
rtx-3090
H100 SXM
h100-sxm
RTX 3090
rtx-3090
H100 PCIe
h100-pcie
RTX 3090
rtx-3090
H100 NVL
h100-nvl
RTX 3090
rtx-3090
A40
a40
RTX 3090
rtx-3090
A100 SXM
a100-sxm
RTX 3090
rtx-3090
A100 PCIe
a100-pcie
RTX 2000 Ada
rtx-2000-ada
RTX A6000
rtx-a6000
RTX 2000 Ada
rtx-2000-ada
RTX A5000
rtx-a5000
RTX 2000 Ada
rtx-2000-ada
RTX A4000
rtx-a4000
RTX 2000 Ada
rtx-2000-ada
RTX 6000 Ada
rtx-6000-ada
RTX 2000 Ada
rtx-2000-ada
RTX 4090
rtx-4090
RTX 2000 Ada
rtx-2000-ada
RTX 3090
rtx-3090
RTX 2000 Ada
rtx-2000-ada
L40S
l40s
RTX 2000 Ada
rtx-2000-ada
L40
l40
RTX 2000 Ada
rtx-2000-ada
L4
l4
RTX 2000 Ada
rtx-2000-ada
H100 SXM
h100-sxm
RTX 2000 Ada
rtx-2000-ada
H100 PCIe
h100-pcie
RTX 2000 Ada
rtx-2000-ada
H100 NVL
h100-nvl
RTX 2000 Ada
rtx-2000-ada
A40
a40
RTX 2000 Ada
rtx-2000-ada
A100 SXM
a100-sxm
RTX 2000 Ada
rtx-2000-ada
A100 PCIe
a100-pcie
L40S
l40s
RTX A6000
rtx-a6000
L40S
l40s
RTX A5000
rtx-a5000
L40S
l40s
RTX A4000
rtx-a4000
L40S
l40s
RTX 6000 Ada
rtx-6000-ada
L40S
l40s
RTX 4090
rtx-4090
L40S
l40s
RTX 3090
rtx-3090
L40S
l40s
RTX 2000 Ada
rtx-2000-ada
L40S
l40s
L40
l40
L40S
l40s
L4
l4
L40S
l40s
H100 SXM
h100-sxm
L40S
l40s
H100 PCIe
h100-pcie
L40S
l40s
H100 NVL
h100-nvl
L40S
l40s
A40
a40
L40S
l40s
A100 SXM
a100-sxm
L40S
l40s
A100 PCIe
a100-pcie
L40
l40
RTX A6000
rtx-a6000
L40
l40
RTX A5000
rtx-a5000
L40
l40
RTX A4000
rtx-a4000
L40
l40
RTX 6000 Ada
rtx-6000-ada
L40
l40
RTX 4090
rtx-4090
L40
l40
RTX 3090
rtx-3090
L40
l40
RTX 2000 Ada
rtx-2000-ada
L40
l40
L40S
l40s
L40
l40
L4
l4
L40
l40
H100 SXM
h100-sxm
L40
l40
H100 PCIe
h100-pcie
L40
l40
H100 NVL
h100-nvl
L40
l40
A40
a40
L40
l40
A100 SXM
a100-sxm
L40
l40
A100 PCIe
a100-pcie
L4
l4
RTX A6000
rtx-a6000
L4
l4
RTX A5000
rtx-a5000
L4
l4
RTX A4000
rtx-a4000
L4
l4
RTX 6000 Ada
rtx-6000-ada
L4
l4
RTX 4090
rtx-4090
L4
l4
RTX 3090
rtx-3090
L4
l4
RTX 2000 Ada
rtx-2000-ada
L4
l4
L40S
l40s
L4
l4
L40
l40
L4
l4
H100 SXM
h100-sxm
L4
l4
H100 PCIe
h100-pcie
L4
l4
H100 NVL
h100-nvl
L4
l4
A40
a40
L4
l4
A100 SXM
a100-sxm
L4
l4
A100 PCIe
a100-pcie
H100 SXM
h100-sxm
RTX A6000
rtx-a6000
H100 SXM
h100-sxm
RTX A5000
rtx-a5000
H100 SXM
h100-sxm
RTX A4000
rtx-a4000
H100 SXM
h100-sxm
RTX 6000 Ada
rtx-6000-ada
H100 SXM
h100-sxm
RTX 4090
rtx-4090
H100 SXM
h100-sxm
RTX 3090
rtx-3090
H100 SXM
h100-sxm
RTX 2000 Ada
rtx-2000-ada
H100 SXM
h100-sxm
L40S
l40s
H100 SXM
h100-sxm
L40
l40
H100 SXM
h100-sxm
L4
l4
H100 SXM
h100-sxm
H100 PCIe
h100-pcie
H100 SXM
h100-sxm
H100 NVL
h100-nvl
H100 SXM
h100-sxm
A40
a40
H100 SXM
h100-sxm
A100 SXM
a100-sxm

B300 vs RTX 6000 Ada

Compare performance across LLMs and image models to find the best GPU for your workload.

Purple glow background

Estimated B300 performance (not measured yet)

We have not run the Runpod benchmark suite on the B300 yet, so the charts below show the RTX 6000 Ada and every other measured GPU without a B300 bar. The table gives our estimate instead: the measured B200 result for each workload, scaled by the one specification that governs it. The B300 (Blackwell Ultra) matches the B200 on memory bandwidth (8 TB/s), dense FP16 tensor throughput (2.25 PFLOPS) and dense FP8 (4.5 PFLOPS), so for BF16 and FP8 inference on 7B to 8B models the defensible estimate is parity with the B200. Where the B300 pulls ahead is 288 GB of HBM3e against 180 GB (1.6x, so larger models and longer contexts fit on one GPU) and 1.5x dense FP4 (15 vs 10 PFLOPS), and neither shows up in these BF16/FP8 tests on small models.

WorkloadRTX 6000 Ada measuredB200 measuredB300 estimatedWhat sets the estimate
Llama 3.1 8B, 1,024 in / 1,024 out, batch 1 (tok/s)52196196Decode is memory-bandwidth-bound at this batch size: 8 TB/s on both parts
Llama 3.1 8B, 1,024 in / 1,024 out, batch 32 (tok/s)9562,6212,621Compute-bound at this batch size: dense FP8/FP16 tensor throughput is unchanged (4.5/2.25 PFLOPS)
Llama 3.1 8B, 1,024 in / 1,024 out, batch 128 (tok/s)1,3973,9283,928Compute-bound at this batch size: dense FP8/FP16 tensor throughput is unchanged (4.5/2.25 PFLOPS)
Qwen2.5 7B, 512 in / 512 out, batch 256 (tok/s)2,42213,76413,764Compute-bound at this batch size: dense FP8/FP16 tensor throughput is unchanged (4.5/2.25 PFLOPS)
Flux.1-dev, 512 px, 10 steps (s/image)2.41.11.1Diffusion is compute-bound in BF16: dense FP16 tensor throughput is unchanged (2.25 PFLOPS)
Flux.1-dev, 1,024 px, 50 steps (s/image)33.110.710.7Diffusion is compute-bound in BF16: dense FP16 tensor throughput is unchanged (2.25 PFLOPS)
SDXL 1.0, 1,024 px, 30 steps (s/image)4.34.44.4Diffusion is compute-bound in BF16: dense FP16 tensor throughput is unchanged (2.25 PFLOPS)

How to read the estimates: expect measured B300 results within roughly 0.95x to 1.10x of the B200 figures. The low end is run-to-run noise; the high end is what the 1,400 W power envelope (against 1,000 W) and Blackwell Ultra's faster attention softmax could add on long-context runs. Values are output tokens per second for LLM rows (higher is better) and seconds per image for diffusion rows (lower is better). Sources: Runpod benchmark service measurements for the B200 and the comparison GPU; NVIDIA Blackwell and Blackwell Ultra specifications. Measured B300 results replace this table as soon as the runs land.

B300
vs.
RTX 6000 Ada
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

LLM inference benchmarks.

Benchmarks were run using vLLM in May 2025 with Runpod GPUs

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
B300

B300

Next-generation data center GPU based on NVIDIA Blackwell Ultra architecture with 288 GB HBM3e memory, delivering 1.5x FP4 performance and 2x attention performance over the B200 for AI reasoning and real-time inference workloads.

RTX 6000 Ada

RTX 6000 Ada

Professional workstation GPU based on Ada Lovelace architecture with 48GB GDDR6 memory and 18,176 CUDA cores for advanced AI workloads.

H100 PCIe

H100 PCIe

High-efficiency LLM processing at 90.98 tok/s.

Image generation benchmarks.

Benchmarks were run using Hugging Face Diffusers in May 2025 on Runpod GPUs.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
H100 SXM

H100 SXM

Unmatched image gen speed with 49.9 images per minute.

H100 NVL

RTX 6000 Ada

AI image processing at 40.3 images per minute.

H100 PCIe

H100 PCIe

Pro-grade performance with 36 images per minute.

Related Comparisons

View All

Real-world GPU
performance in action.

See how teams optimize cost and performance with the right GPU for their workloads.
How TOOL Scales Big AI Ideas on Runpod

"All of these projects, the renders for AMD, the Coca-Cola builds, that has to do with scalability. If we can't scale, we can't deliver. Runpod makes that possible."

How Aneta Handles Bursty GPU Workloads Without Overcommitting

"Runpod has changed the way we ship because we no longer have to wonder if we have access to GPUs. We've saved probably 90% on our infrastructure bill, mainly because we can use bursty compute whenever we need it."

How Gendo uses Runpod Serverless for Architectural Visualization

"Runpod has allowed the team to focus more on the features that are core to our product and that are within our skill set, rather than spending time focusing on infrastructure, which can sometimes be a bit of a distraction.”

How Civitai Trains 800K Monthly LoRAs in Production on Runpod

"Runpod helped us scale the part of our platform that drives creation. That’s what fuels the rest, image generation, sharing, remixing. It starts with training."

How Scatter Lab Powers 1,000+ Inference Requests per Second with Runpod

"Runpod allowed us to reliably handle scaling from zero to over 1,000 requests per second in our live application."

How InstaHeadshots Scales AI-Generated Portraits with Runpod

"Runpod has allowed us to focus entirely on growth and product development without us having to worry about the GPU infrastructure at all."

How KRNL AI scaled to 10K+ concurrent users while cutting infra costs 65%.

"We could stop worrying about infrastructure and go back to building. That’s the real win.”

How Coframe scaled to 100s of GPUs instantly to handle a viral Product Hunt launch.

“The main value proposition for us was the flexibility Runpod offered. We were able to scale up effortlessly to meet the demand at launch.”

How Glam Labs Powers Viral AI Video Effects with Runpod

"After migration, we were able to cut down our server costs from thousands of dollars per day to only hundreds."

How Segmind Scaled GenAI Workloads 10x Without Scaling Costs

Runpod’s scalable GPU infrastructure gave us the flexibility we needed to match customer traffic and model complexity, without overpaying for idle resources.

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.

Star field background