News icon

Kimi K3 is now available on Runpod

Global Volumes are now in beta for Serverless endpoints

Attach a Global Volume to a Runpod Serverless endpoint and your workers start in any data center with your GPU types. Store model weights once, read anywhere.

Global Volumes are now in beta for Serverless endpoints

Storage can pin your Serverless endpoint to one region. When your model weights live on a network volume tied to one data center, your workers have to start there too. 

Global Volumes, now in beta for Runpod Serverless, remove that constraint. Store your model weights once and let workers start in any data center with your selected GPU types.

Today we're launching Global Volumes for Runpod Serverless in beta. Global Volumes are elastic, region-independent storage, and you can now attach them to Serverless endpoints as well as Pods. Store a model once, and your endpoint's workers can rapidly cold start in any data center that has your GPU types available. No need to know your storage needs up front – pay as you go without setting a fixed volume size.

Diagram: model weights are copied once from a Pod into a Global Volume. Serverless workers on one endpoint mount that same volume at /runpod-volume in several data centers. When data center A has no free capacity, the scheduler places a new worker in data center C, and it reads the same files.

GLOBAL VOLUMES FOR SERVERLESS
One copy of your model, workers in any data center
Copy weights in once from a Pod. Every worker on the endpoint mounts the same volume, wherever your GPU types are available.
Global Volumes for Serverless SERVERLESS ENDPOINT Pod copy models in Global Volume not tied to a data center Worker DC A /runpod-volume Worker DC B /runpod-volume Worker DC C /runpod-volume WRITE ONCE ONE COPY DC A AT CAPACITY PLACED IN DC C
NO PER-REGION COPIES MOUNTS AT /runpod-volume BETA

Global volumes are optimized for use cases like storing model weights. Train in one data center and serve in another. Write a base model, low-rank adaptation (LoRA) adapter, tokenizer, or config once, and use the same global storage. For maximum write throughput and ultra low latency workloads, high performance Network Volumes are a better choice but the tradeoff is they lock your GPU workers to a region.

How it works

Global Volumes are not tied to a data center, so the scheduler places the endpoint's workers wherever your selected GPU types are available.

Runpod mounts the volume into each worker at /runpod-volume, the same path network volumes use on Serverless. If your handler already loads models from /runpod-volume, it runs as is, with no code change or image rebuild.

Reads reflect the state of the volume at mount time. A running worker will not pick up a file you upload after it started, so roll out new weights by replacing workers, not by writing to the volume underneath them.

Use cases and trade-offs

Global Volumes are designed for data that is written infrequently and read across deployments: model weights, LoRAs, tokenizers, configuration and similar artifacts. One copy serves every region, so a model update doesn't need to be copied to a volume in each data center.

Capacity grows with what you store, so there's nothing to size up front or resize later. Billing is usage-based: you pay for data stored plus the requests made against it. Writes, creates, and lists are billed at $0.005 per 1,000 requests; reads and metadata lookups at $0.0005 per 1,000 requests. A single file system call can issue several requests, so test a workload at small scale before pointing a busy endpoint at a volume.

Global Volumes are backed by object storage and are not a replacement for high-performance regional storage. Workloads that depend on full POSIX semantics, concurrent writes from multiple workers to the same file, high write throughput, or the lowest possible read latency should use a network volume instead.

Keep your Python environment in the container image and point your model directory or Hugging Face cache (HF_HOME) at /runpod-volume. For a ComfyUI worker on Runpod Serverless, that means ComfyUI and its custom nodes live in the image, and extra_model_paths.yaml points the checkpoint and LoRA directories at /runpod-volume. 

# Dockerfile
ENV HF_HOME=/runpod-volume/hf

# extra_model_paths.yaml
runpod_volume:
  base_path: /runpod-volume/
  checkpoints: models/checkpoints/
  loras: models/loras/

With this config, checkpoints go in /runpod-volume/models/checkpoints/ and LoRAs in /runpod-volume/models/loras/ on the Global Volume. Applications that install onto the volume, such as a full ComfyUI setup or a virtual environment, depend on file permissions the volume doesn't support and will fail or run slowly.

Global Volumes compared with network volumes

Global Volume Network volume
Data center binding None Pinned to one data center
Mount path on Serverless /runpod-volume /runpod-volume
Best for Read-heavy files: model weights, LoRAs, tokenizers, configs Write-heavy work: training, checkpointing, high-throughput workloads
File system Object-backed; no file locking, atomic rename or hard links Full file system semantics on NVMe SSDs, 200–400 MB/s typical

Because the mount path is the same, moving an endpoint from one to the other doesn't change your handler code.

Getting started

If your endpoint's network volume is the reason it can only run in one data center, move its models to a Global Volume.

The beta is available to all users. To create a Global Volume, go to Storage, click + New volume, and set the storage type to Global volume. No data center selection is required. To add models, attach the volume to a Pod and copy files into it. On a Pod, a Global Volume mounts at /workspace, or at /workspace-global when a network volume is attached alongside it, so you can copy models from an existing network volume to the Global Volume on the same Pod.

When you create and manage an endpoint, pick Global Volume from the Advanced section. The endpoint overview shows the attached volume's size and object count. 

Four limits apply during the beta:

An endpoint can have at most one Global Volume.

  • CPU workers aren't supported yet.
  • An endpoint can use a network volume or a Global Volume, not both, because both mount at /runpod-volume. To switch, remove one and select the other.
  • If your account balance reaches $0, the volume is flagged and permanently deleted after 15 days. Turn on low balance notifications and automatic recharges.

Setup details are in the Global Volumes for Serverless docs. During the beta, use the feedback link in the console to reach the engineering team. 

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.

Star field background