News icon

Kimi K3 is now available on Runpod

NVIDIA B300 capacity, held for your workload

Production B300 capacity without the 18-month procurement cycle. Talk to us about held capacity, a committed per-GPU-hour rate, and the regions you need.

Reserved B300 capacity, multi-node Clusters, and volume pricing.

Slurm orchestration illustration
Purple glow background

288 GB of HBM3e per GPU

50% more memory than the B200, from 12-high HBM3e stacks. A large model shards across fewer devices.

Dynamic node management illustration
Purple glow background

8 TB/s of memory bandwidth

Bandwidth is what feeds the tensor cores, and 15 petaFLOPS of dense NVFP4 compute sits behind it.

Cluster monitoring illustration
Purple glow background

1.8 TB/s NVLink 5 per GPU

Multi-node B300 Clusters over InfiniBand, deployed in minutes rather than days.

Trusted by top engineers at the world's leading companies.
Purple glow background

Why teams run B300 workloads on Runpod

Reserved capacity you can hold, the whole path from training run to inference endpoint, and a migration measured in days.

Training and fine-tuning large models

288 GB per GPU holds a model and its optimizer state on fewer devices, which cuts the tensor parallelism you have to configure and the cross-GPU communication that comes with it.

Long-context and reasoning inference

Reasoning models emit far more tokens per request than chat models, and the KV cache is what fills up first. More HBM per GPU means more concurrent requests before you start sharding.

Multi-node Clusters

When one node runs out, Clusters give you InfiniBand-connected B300 nodes deployed in minutes rather than days. Ask us about node counts and topology for your run.

Capacity you can hold

Reserved capacity and Savings Plans lock a per-GPU-hour rate to a specific GPU type across 31 global regions, so you get production AI infrastructure without the 18-month procurement cycle.

One platform, the whole lifecycle

Pods for the training run, Serverless for the inference endpoint it becomes, Clusters when it outgrows a single node. Same account and no replatforming between stages.

Migration measured in days

Gendo moved their entire GPU backend off AWS in a day with the same containers. Glam Labs finished their migration in a day and cut server costs 90%.

What is the NVIDIA B300?

The B300 is NVIDIA's Blackwell Ultra data center GPU. NVIDIA's published specifications put it at 288 GB of HBM3e, 8 TB/s of memory bandwidth, and 15 petaFLOPS of dense NVFP4 compute, with 1.8 TB/s of NVLink 5 bandwidth per GPU. Runpod runs B300 as on-demand Pods, as Serverless workers, and as multi-node Clusters.

Need reserved or multi-node B300 capacity? Talk to our team.

"The Runpod team has clearly prioritized the developer experience to create an elegant solution that enables individuals to rapidly develop custom AI apps or integrations while also paving the way for organizations to truly deliver on the promise of AI."

Amjad Masad

"Runpod is the only place I can deploy high-end GPU models instantly. No sales calls, no rate limits, no nonsense."

Daniel Chang

“The main value proposition for us was the flexibility Runpod offered. We were able to scale up effortlessly to meet the demand at launch.”

Josh Payne

“Runpod helped us scale the part of our platform that drives creation. That’s what fuels the rest. Image generation, sharing, remixing. It starts with training.”

Matty Shimura

Runpod reservation savings illustration

Reserve B300 capacity before the run starts

A Savings Plan locks a per-GPU-hour rate for a specific GPU type, so you know what a long training run costs before you commit to it. Compute on Runpod runs up to 90% lower than traditional cloud providers, on structural pricing rather than a discount with an expiry date.

B300 questions, answered

What the hardware does, what you can reserve, and what happens when you talk to us.

The B300 is NVIDIA's Blackwell Ultra data center GPU. NVIDIA's published specifications put it at 288 GB of HBM3e, 8 TB/s of memory bandwidth, and 15 petaFLOPS of dense NVFP4 compute, with 1.8 TB/s of NVLink 5 bandwidth per GPU. On Runpod you can run it as an on-demand Pod, as a Serverless worker, or as part of a multi-node Cluster.
Same Blackwell architecture with more memory stacked on it. The B300 uses 12-high HBM3e stacks where the B200 uses 8-high, which is where the 288 GB and the extra low-precision throughput come from. If your constraint is memory per GPU, that is the difference that matters.
Yes. Reserved capacity and Savings Plans lock a per-GPU-hour rate for a specific GPU type, and multi-node B300 Clusters are available on request. Talk to sales and we will scope node count, region, and term with you.
No. On-demand B300 Pods are self-serve and billed by the second. A contract is for teams that want held capacity, a committed rate, post-paid billing, SSO, or a technical account manager.
Current on-demand rates are on the pricing page and bill by the second. Reserved and multi-node pricing depends on term, region, and node count, which is what the sales conversation is for.
We operate 31 global regions and B300 availability varies by datacenter, so ask us for current capacity in the regions you need.

OpenAI, Perplexity, Replit, Cursor, Wix and Zillow build on Runpod.

Engineered for teams building the future.

Wix logo
Otovo logo
Scatter Lab logo
Abzu logo
Aneta logo
Perplexity logo
Replit logo
Civitai logo

Tell us what you are training

Send us the model, the scale, and the timeline. We will come back with the B300 capacity we can hold, the regions it sits in, and what it costs at your volume.

Star field background