The NVIDIA B300 is the Blackwell Ultra refresh of the Blackwell architecture, and the largest-memory GPU available on Runpod. It carries 288 GB of HBM3E, 8 TB/s of memory bandwidth, and 15 petaFLOPS of dense NVFP4 compute – 50% more memory and 50% more low-precision throughput than the B200, from the same 208-billion-transistor dual-die design.
Runpod offers on-demand B300 instances from $6.94/hr on Community Cloud, billed by the second. This guide covers the full B300 specs, what Blackwell Ultra changes versus Blackwell, how it compares to the B200 and H200, and when the extra memory is worth paying for.
How much VRAM does the B300 have?
The NVIDIA B300 has 288 GB of HBM3E memory with 8 TB/s of bandwidth. That is 3.6x the H100's 80 GB and 50% more than the B200, from eight 12-high HBM stacks running across sixteen 512-bit controllers – an 8,192-bit total interface.
Across the GPUs on Runpod:
- B300: 288 GB HBM3E, 8 TB/s
- B200: 180 GB HBM3e, ~8 TB/s
- H200 SXM: 141 GB HBM3e, ~4.8 TB/s
- H100 SXM: 80 GB HBM3, ~3.35 TB/s
- A100 80GB: 80 GB HBM2e, ~2 TB/s
Note that bandwidth is unchanged from the B200 at 8 TB/s. Blackwell Ultra spends its gains on capacity and compute, not on feeding the memory faster. That matters for how you choose between them: if your bottleneck is bandwidth, the B300 will not fix it. If your bottleneck is capacity – a model that will not fit, or a KV cache you keep having to evict – that is exactly what the extra 108 GB buys.
NVIDIA B300 specs
- Architecture: Blackwell Ultra, dual-die, TSMC 4NP
- Transistors: 208 billion
- Die-to-die interconnect: NV-HBI at 10 TB/s
- VRAM: 288 GB HBM3E
- Memory bandwidth: 8 TB/s
- SMs: 160, across eight GPCs
- CUDA cores: 20,480
- Tensor Cores: 640, fifth generation with second-gen Transformer Engine
- NVFP4: 15 petaFLOPS dense, 20 petaFLOPS sparse
- FP8: 5 petaFLOPS dense, 10 petaFLOPS sparse
- Attention acceleration: 10.7 TeraExponentials/s (SFU EX2), roughly 2x Blackwell
- NVLink: fifth generation, 1.8 TB/s bidirectional per GPU
- PCIe: Gen6 x16, 256 GB/s bidirectional
- NVLink-C2C: 900 GB/s to a Grace CPU
- MIG: supported – two instances at 140 GB, four at 70 GB, or seven at 34 GB
- Decompression engine: 800 GB/s
- Max power (TGP): up to 1,400W
B300 vs B200: what Blackwell Ultra changes
The B300 and B200 share a transistor count, a process node and a memory bandwidth figure. Three things separate them:
- Memory: 288 GB against 180 GB. The 12-high HBM stacks are the whole reason the B300 exists.
- NVFP4 compute: 15 petaFLOPS dense against Blackwell's 10 – a 50% increase. Sparse throughput is unchanged at 20 petaFLOPS, so the gain is specific to dense inference.
- Attention: SFU throughput for the exponentials inside softmax is roughly doubled, at 10.7 TeraExponentials/s against 5. On long-context reasoning workloads, softmax is often where the latency actually goes, and this is the least-discussed and most useful difference between the two cards.
FP8 throughput is identical on both at 5 petaFLOPS dense. If your stack has not moved to NVFP4, the B300's compute advantage does not reach you – you are paying for the memory alone.
B300 vs H200 and H100
- Against the H200: 288 GB versus 141 GB, and 8 TB/s versus 4.8 TB/s. The H200 is a memory upgrade on the Hopper die; the B300 is two generations of architecture ahead of it, with NVFP4 and fifth-gen NVLink.
- Against the H100: 3.6x the memory and 2.4x the bandwidth. NVIDIA puts Blackwell Ultra's NVFP4 throughput at 7.5x Hopper's FP8.
- On price: a B300 is $6.94/hr on Community Cloud against $2.69/hr for an H100 SXM on the same tier – roughly 2.6x the hourly rate. Whether that pays depends entirely on whether you can saturate it.
B300 price on Runpod
- Community Cloud: $6.94/hr
- Secure Cloud: $7.39/hr
- Serverless: $9.98/hr, billed only while requests run
Each pod instance includes 288 GB VRAM, 251 GB RAM and 32 vCPUs. Everything is billed by the second, so a two-hour evaluation run costs under $14 rather than a procurement cycle. For reserved or multi-node B300 capacity, contact the Runpod sales team. See current availability on the Runpod pricing page.
When to use a B300
The B300 is the right call when:
- Your model does not fit in 180 GB at the precision you need
- You are serving long-context or reasoning models where KV cache, not weights, is what fills the card
- Your inference stack runs NVFP4, so the 50% dense compute gain is reachable
- Consolidating two B200s onto one B300 removes a tensor-parallel split and the interconnect overhead that comes with it
Something smaller is the better buy when:
- Your model fits comfortably in 180 GB – the B200 starts at $5.89/hr on Secure Cloud and does the same work for less
- You are bandwidth-bound rather than capacity-bound, since both cards sit at 8 TB/s
- You are running FP8 and have no NVFP4 path, which erases most of the compute difference
- Your workload fits in 141 GB, where the H200 at $3.59/hr on Community Cloud is roughly half the price
Rent a B300 on Runpod
Blackwell Ultra keeps full CUDA compatibility, and vLLM, SGLang and TensorRT-LLM all support it natively with NVFP4 kernels – moving a working stack across is a version bump, not a rewrite.
- Pods: from $6.94/hr on Community Cloud, $7.39/hr on Secure Cloud, billed per second
- Serverless: B300 workers at $9.98/hr that scale to zero when idle, so you pay nothing between requests
- Network volumes: persistent storage for weights and checkpoints, from $0.07/GB/mo standard and $0.14/GB/mo high-performance. One volume can be shared across instances, so you stage a large model once
- Templates: pre-built PyTorch and vLLM environments with CUDA already configured
- Flexibility: move between B300, B200, H200, H100 and the rest without hardware lock-in
B300 FAQs
How much VRAM does the B300 have?
288 GB of HBM3E with 8 TB/s of bandwidth – 50% more capacity than the B200 and 3.6x the H100.
How much does it cost to rent a B300?
On Runpod, $6.94/hr on Community Cloud and $7.39/hr on Secure Cloud, billed by the second. Serverless B300 workers are $9.98/hr and cost nothing while idle. Each pod includes 288 GB VRAM, 251 GB RAM and 32 vCPUs.
What is the difference between the B300 and B200?
Same transistor count, same process, same 8 TB/s bandwidth. The B300 adds 108 GB of memory, 50% more dense NVFP4 compute, and roughly double the attention-layer throughput. If your model fits in 180 GB and you are not running NVFP4, the B200 is the better value.
What is Blackwell Ultra?
Blackwell Ultra is the second generation of NVIDIA's Blackwell architecture. It keeps the dual-die design and 208 billion transistors, and increases HBM capacity to 288 GB, dense NVFP4 throughput to 15 petaFLOPS, and softmax execution speed to 10.7 TeraExponentials/s. The B300 is the GPU; GB300 refers to the Grace Blackwell Ultra rack-scale systems built from it.
What is NVFP4?
NVFP4 is a 4-bit floating-point format introduced with Blackwell. It applies two levels of scaling – an FP8 micro-block scale across 16 values plus a tensor-level FP32 scale – which keeps accuracy close to FP8, typically within about 1%, while cutting memory footprint by roughly 1.8x against FP8 and 3.5x against FP16.
What is the difference between the B300 and GB300 NVL72?
The B300 is a single GPU. GB300 NVL72 is a liquid-cooled rack containing 36 Grace Blackwell Ultra Superchips – 72 B300 GPUs alongside 36 Grace CPUs – reaching 1.1 exaFLOPS of dense FP4. Runpod offers individual B300 instances, giving you the same silicon without buying a rack.
