News icon

Kimi K3 is now available on Runpod

Rent NVIDIA B300 GPUs in the Cloud on Runpod

Rent NVIDIA B300 GPUs on Runpod from $6.94/hr on Community Cloud and $7.89/hr on Secure Cloud, billed by the second, in pods of 1 to 8 GPUs. The B300 is the Blackwell Ultra GPU with 288 GB of HBM3e per card, which is 108 GB more than the B200 and 3.6x the H100's 80 GB. Rent cloud GPUs on Runpod and a B300 pod is running in under a minute, with no quote, no contract, and no minimum term.

Why rent the NVIDIA B300

The B300 is built for models that do not fit anywhere else. It keeps the B200's dual-die design, 208 billion transistors and 8 TB/s of memory bandwidth, then raises HBM capacity to 288 GB and dense NVFP4 throughput to 15 petaFLOPS. If your model fits in 180 GB at the precision you need, the B200 at $5.98/hr does the same work for less. If it does not, the B300 is the card that keeps it on one GPU instead of two.

Benefits

  • 288 GB of HBM3e per GPU
    The largest memory pool on any single GPU Runpod rents. An 8xB300 pod holds 2,304 GB of GPU memory in one NVLink domain, enough to serve a 2.8-trillion-parameter model at 4-bit precision on a single node. In August 2026 Runpod's Brendan McKeag deployed Kimi K3 this way; its 1,560.94 GB checkpoint does not load on an 8xH100 node (640 GB) or an 8xB200 node (1,536 GB).
  • 15 petaFLOPS of dense NVFP4
    Blackwell Ultra adds 50% more dense FP4 compute than the B200 and roughly doubles attention-layer throughput, which is where long-context inference spends its time. NVIDIA reports NVFP4 accuracy within about 1% of FP8 with a 1.8x smaller memory footprint.
  • Fifth-generation NVLink at 1.8 TB/s
    Every B300 pod on Runpod is an SXM configuration with NVLink between GPUs, so tensor-parallel and pipeline-parallel jobs scale across the pod without a PCIe bottleneck. See how it compares with the previous generation in B300 vs H200 and B300 vs H100 SXM.
  • Per-second billing with no commitment
    A two-hour evaluation run on one B300 costs under $14 on Community Cloud. You pay for the seconds the pod is running, so a failed job or an early stop costs exactly what it used. Current rates are on the Runpod pricing page.
  • Multi-node when you need it
    For training or inference that outgrows one 8-GPU node, Runpod Clusters connect multiple B300 nodes with high-speed interconnect. Clusters are reserved capacity; contact sales for B300 cluster pricing.

B300 price on Runpod

ProductB300 priceBilling
Pods, Community Cloud$6.94/hrPer second, on demand
Pods, Secure Cloud$7.89/hrPer second, on demand
Clusters (multi-node B300)Contact salesReserved

Each B300 pod includes 288 GB of VRAM, 251 GB of system RAM and 32 vCPUs per GPU. Persistent storage for weights and checkpoints is $0.07/GB/mo under 1 TB and $0.05/GB/mo above it on standard network volumes, or $0.14/GB/mo on high-performance storage. Prices checked September 12, 2026 against the Runpod pricing page; the B300 model page shows live availability by region.

Specifications

FeatureValue
ArchitectureBlackwell Ultra
DesignDual-die, 208 billion transistors, same process as the B200
Form factorSXM module (HGX B300 baseboard)
Memory capacity288 GB HBM3e
Memory bandwidth8 TB/s
NVFP4 tensor performance15 PFLOPS dense
FP8 tensor performance5 PFLOPS dense, 10 PFLOPS with sparsity
Attention throughputAbout 2x the B200 (10.7 TeraExponentials/s softmax)
GPU-to-GPU interconnectFifth-generation NVLink, 1.8 TB/s
Per-GPU pod allocation on Runpod251 GB RAM, 32 vCPUs
Pod sizes on Runpod1 to 8 GPUs; 2,304 GB HBM3e per 8-GPU pod

For the full spec sheet, what changed from the B200, and how the B300 relates to the GB300 NVL72 rack, read the NVIDIA B300 guide.

FAQ

How much does it cost to rent a B300 on Runpod?

$6.94/hr on Community Cloud and $7.89/hr on Secure Cloud, billed per second with no minimum. A B300 costs about 2.6x an H100 SXM per hour on the same Community Cloud tier, and about 16% more than a B200.

Is there a minimum rental duration?

No. Runpod bills by the second and you can stop a pod at any time. Reserved pricing for longer commitments is available through Clusters; contact sales.

When is the B300 worth it over the B200 or H200?

When capacity is the bottleneck. The B300 and B200 share the same 8 TB/s bandwidth, so a bandwidth-bound workload that fits in 180 GB runs about as fast on the B200 for $5.98/hr. A workload that fits in 141 GB runs on the H200 for $3.59/hr. The B300 earns its price when the model, its KV cache, or a training state does not fit in 180 GB at the precision you need.

How many B300s do I need?

Count memory first. At 4-bit precision, roughly 0.55 GB per billion parameters plus KV cache: a 400B model fits on one B300, a 1T model wants four, and Kimi K3 at 2.8T needs all eight. For training, plan on 16 to 20 bytes per parameter for weights, gradients and optimizer state before activations.

Can I run multi-node B300 jobs?

Yes, through Runpod Clusters, which connect multiple 8xB300 nodes for distributed training and inference. Clusters are reserved capacity with pricing through sales, unlike pods, which are on demand.

What software comes with a B300 pod?

Pods start from a container image you choose. Runpod's official PyTorch templates ship with CUDA and the current NVIDIA driver, and vLLM and SGLang templates are in the Runpod Hub. Blackwell Ultra needs a CUDA 12.9 or newer base image.

Where are B300 pods available?

B300 is the newest Blackwell generation and supply moves week to week. The B300 model page shows live stock by region, and the console deploy page filters to data centers with B300 capacity right now.

Is my data isolated?

Secure Cloud pods run in T3 and T4 data centers under Runpod's SOC 2 Type II program, with dedicated GPU allocation. Community Cloud pods run on vetted hosts at a lower price. Both bill per second; pick Secure Cloud for regulated workloads and Community Cloud for evaluation and cost-sensitive jobs.

Author profile: The Runpod Team

Purple glow background

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.

Star field background