
RTX A4000
Professional single-slot GPU based on Ampere architecture with 16GB GDDR6 memory and 6,144 CUDA cores for AI workloads, machine learning, and compact workstation builds.
GPU Benchmarks
Compare performance across LLMs and image models to find the best GPU for your workload.
Runpod benchmark measurements from September 2026, using vLLM.

Professional single-slot GPU based on Ampere architecture with 16GB GDDR6 memory and 6,144 CUDA cores for AI workloads, machine learning, and compact workstation builds.
.avif)
Dual-GPU data center accelerator based on Hopper architecture with 188GB combined HBM3 memory (94GB per GPU) designed specifically for LLM inference and deployment.

High-efficiency LLM processing at 90.98 tok/s.
Runpod benchmark measurements from September 2026, using Hugging Face Diffusers.

Unmatched image gen speed with 49.9 images per minute.

AI image processing at 40.3 images per minute.

Pro-grade performance with 36 images per minute.
Case Studies