News icon

Kimi K3 is now available on Runpod

Rent AMD MI300X GPUs from $2.39/hr

Data center accelerator built on AMD CDNA 3 architecture with 192 GB of HBM3 memory and 5.3 TB/s of bandwidth, for large-model inference and training on the ROCm software stack.

MI300X

Powering the next generation of AI & high-performance computing.

Engineered for large-scale AI training, deep learning, and high-performance workloads, delivering unprecedented compute power and efficiency.

AMD CDNA 3 Architecture

Chiplet design with 304 compute units and 1,216 matrix cores across 153 billion transistors, built for generative AI and HPC.

192 GB HBM3 Memory

The largest memory capacity of any accelerator we offer, at 5.3 TB/s with full-chip ECC and 256 MB of Infinity Cache.

ROCm Open Software Stack

Open-source toolchain with PyTorch and vLLM support, so most inference workloads move across without kernel-level rewrites.

Infinity Fabric Interconnect

Eight Infinity Fabric links at up to 128 GB/s each connect accelerators directly, for multi-GPU work across up to eight cards.

Why rent the MI300X instead of buying?

The most memory per accelerator available here

At 192 GB of HBM3, the MI300X carries more than twice the memory of an H100 SXM and roughly a third more than an H200. For inference on very large models, memory capacity is the constraint that decides whether a model runs on one accelerator or has to be split across several, and splitting costs you latency and complexity. Bandwidth is 5.3 TB/s against the H100 SXM's 3.35 TB/s, so the extra capacity is not paid for with slower access.

Hardware you cannot put in a workstation

The MI300X ships as a passively cooled OAM module drawing up to 750W, on a 54V UBB power delivery system. It is not a card you drop into a chassis. Renting removes the power, cooling and integration problem entirely, and lets you find out whether your stack runs well on ROCm before committing to anything.

Test the ROCm question cheaply

The real question with AMD is rarely the hardware. It is whether your stack runs on ROCm rather than CUDA. Most mainstream inference paths, including PyTorch and vLLM, are supported, but custom kernels and niche libraries are where projects find friction. Per-second billing with no minimum means you can answer that question with a real workload in an afternoon rather than a procurement cycle.

Key specs at a glance.

Performance benchmarks that push AI, ML, and HPC workloads further.

Memory Capacity

192

GB HBM3

Memory Bandwidth

5.3

TB/s

FP8 Performance

2.61

PFLOPS

Popular use cases.

Designed for demanding workloads. Learn if this GPU fits your needs.

Inference workload illustration

Inference

Serve inference for image, text, and audio generation at any scale.

Fine-tuning workload illustration

Fine-tuning

Train custom models on your specific datasets.

AI agents workload illustration

Agents

Build intelligent agent-based systems and workflows.

Compute-heavy workload illustration

Compute-heavy tasks

Run compute-heavy workloads like rendering and simulations.

Ready for your most demanding workloads.

Essential technical specifications to help you choose the right GPU for your workload.

Specification
Details
Great for...
Memory Capacity
192 GB HBM3
Holding very large models on a single accelerator without sharding, and serving long-context workloads where KV cache is the real constraint.
Memory Bandwidth
5.3 TB/s
Feeding weights to the compute units on memory-bound inference, where bandwidth rather than raw arithmetic sets the ceiling.
FP8 Performance
2.61 PFLOPS
Quantized production inference at high throughput, with 5.22 PFLOPS available on workloads that can use structured sparsity.
SpecificationDetailsGreat for...
ArchitectureAMD CDNA 3Large-model inference and HPC on the ROCm stack, where memory capacity matters more than CUDA compatibility
Manufacturing ProcessTSMC 5nm / 6nm FinFETN/A
Transistors153 billionN/A
Compute Units304N/A
Stream Processors19,456N/A
Matrix Cores1,216Accelerating the matrix arithmetic behind transformer training and inference
GPU Memory192 GB HBM3Running very large models on a single accelerator without sharding, and long-context inference where KV cache dominates
Memory Bandwidth5.3 TB/sMemory-bound inference, where bandwidth rather than arithmetic sets the ceiling
Memory Interface8,192-bitN/A
Infinity Cache256 MBN/A
Memory ECCYes, full-chipLong-running jobs where a silent bit flip would invalidate the result
Peak Engine Clock2,100 MHzN/A
Board Power (TBP)750W peakRack configurations with the power and cooling to match
Form FactorOAM module, passive coolingN/A
System InterfacePCIe 5.0 x16High-speed host-to-accelerator transfers for large dataset pipelines
Infinity Fabric Links8 links, up to 128 GB/s eachDirect accelerator-to-accelerator communication across up to eight cards
FP64 Performance81.7 TFLOPS vector, 163.4 TFLOPS matrixHigh-precision scientific computing and simulation
FP32 Performance163.4 TFLOPSStandard-precision training and inference
TF32 Matrix653.7 TFLOPSAccelerated training with near-FP32 accuracy
FP16 / BF161,307.4 TFLOPSMixed-precision training and high-throughput inference
FP82,614.9 TFLOPSMaximum throughput on quantized production models
INT82.6 POPSQuantized inference at scale
SR-IOVYesHardware-level partitioning for multi-tenant serving

"The Runpod team has clearly prioritized the developer experience to create an elegant solution that enables individuals to rapidly develop custom AI apps or integrations while also paving the way for organizations to truly deliver on the promise of AI."

Amjad Masad

"Runpod is the only place I can deploy high-end GPU models instantly. No sales calls, no rate limits, no nonsense."

Daniel Chang

“The main value proposition for us was the flexibility Runpod offered. We were able to scale up effortlessly to meet the demand at launch.”

Josh Payne

“Runpod helped us scale the part of our platform that drives creation. That’s what fuels the rest. Image generation, sharing, remixing. It starts with training.”

Matty Shimura

Powerful GPUs. Globally available. Reliability you can trust.

30+ GPUs, 31 regions, instant scale. Fine-tune or go full Skynet. We’ve got you.

Community Cloud
$/hr
Secure Cloud
$2.39/hr
Unique GPU Models
Community Cloud
25
Secure Cloud
19
Global Regions
Community Cloud
17
Secure Cloud
14
Network Storage
Community Cloud
Secure Cloud
Enterprise-Grade Reliability
Community Cloud
Secure Cloud
Savings Plans
Community Cloud
Secure Cloud
24/7 Support
Community Cloud
Secure Cloud
Delightful Dev Experience
Community Cloud
Secure Cloud

Questions? Answers.

What are the current hourly rates for renting an MI300X on Runpod?


The MI300X is available on Secure Cloud. For the most current pricing, see the Runpod pricing page.

Will my code run on an MI300X?


It depends on your stack rather than your model. The MI300X runs AMD's ROCm software rather than CUDA. Mainstream inference paths including PyTorch and vLLM are supported, and most workloads that use them move across without changes. Custom CUDA kernels and libraries with no ROCm equivalent are where projects hit friction, so the practical advice is to test your own workload rather than assume either way.

How does the MI300X compare to an H100 or H200?


On memory it is well ahead: 192 GB of HBM3 at 5.3 TB/s, against 80 GB at 3.35 TB/s on an H100 SXM and 141 GB at 4.8 TB/s on an H200. That matters most for large-model inference and long-context work, where capacity decides whether a model fits on one accelerator. The trade is software: the NVIDIA cards run CUDA, which almost everything targets first. Choose on whether your stack is portable, not on the specification sheet alone.

How many MI300X GPUs can I deploy?


Configurations of up to eight accelerators are available, connected by Infinity Fabric links at up to 128 GB/s each. See the Runpod pricing page for current multi-GPU options.

Is the MI300X suitable for sensitive or regulated workloads?


Runpod is SOC 2 Type II certified and HIPAA and GDPR compliant, with SOC 2 reports, BAAs, and DPAs available for security review through the Runpod Trust Center. Secure Cloud adds network isolation for workloads with stricter compliance needs. Coverage can vary by region and deployment model, so check requirements for your workload. See runpod.io/legal/compliance for current certification details.

10,100,100,100

Requests since launch & 1M+ developers worldwide

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.

Star field background