News icon

Kimi K3 is now available on Runpod

All your AI compute, one control plane.

Bring the compute you already own or rent under the same console, CLI, and APIs your developers trust on Runpod, with your hardware running first and Runpod cloud there for overflow when you need it.

The same layer that runs the Runpod cloud, now on your hardware.

Already build, train, and ship on Runpod.

Runpod orchestrates it and never takes ownership.

Workloads land on your compute, with overflow when you need it.

The problem

When compute is scattered, utilization is a guess.

GPUs spread across on-prem, hyperscalers, and vendor clouds, with no single way to see, allocate, or govern them. Teams over-provision and still wait in line.

How it works

From your hardware to your developers, in three steps.

No replatforming. The hardware you already run becomes dedicated compute your team deploys to from the Runpod console.

Step 1

Install the host agent

Run one install command on hardware you own. The host agent connects outbound over an encrypted overlay, no inbound ports.

Step 2

Verify into your dedicated capacity

Runpod verifies the hardware, then promotes it into capacity scoped to your org, visible only to your organization.

Step 3

Launch on your hardware, first

Your team launches pods from the same console and APIs. Workloads run on your hardware first, with overflow to Runpod cloud when you need it.

The platform

One control plane for the compute you own and the compute you rent.

Three commitments, one system: your estate managed as one, the developer experience your team already knows, and a boundary that stays yours.

All your AI compute, one control plane

Manage the compute you own and the Runpod cloud as one estate, owned hardware first and cloud as overflow.

Consistent developer experience

Give your team the Runpod console, CLI, and APIs on the hardware you own. SSH still works; the queue scripts and access tickets stop being the only way in.

Your hardware, your data, your boundary

You own the hardware and the data on it. Runpod orchestrates it, and it stays visible only to your org.

Who it's for

One platform, tuned to how your team runs.

Modern self-service on the cluster you already run.

Researchers get the Runpod console on the institution's GPUs. The center keeps the boundary, and Runpod cloud absorbs overflow when the queue is full.

Every GPU you own or rent, managed as one estate.

Bring on-prem, colo, and cloud capacity under one control plane. Owned hardware runs first, and Runpod cloud is the governed overflow.

Get full value from the capacity you already pay for.

Put your reserved or owned cluster to work through Runpod, and burst to the cloud for spikes, without standing up a platform team.

Security and data

Your hardware and your data stay yours.

A posture you can take straight into a security review, with a direct line to our engineers for everything underneath it.

Outbound-only host agent

Connects out over an encrypted overlay, with no inbound ports opened.

Private to your org

Everything you onboard stays visible only to your organization.

Your data stays put

Workloads run on your hardware; data on local NVMe you own.

No lock-in

The host agent uninstalls and your workloads are standard containers.

Why it's different

Orchestrate the compute that's already yours.

Without an orchestration layer

With Runpod Hybrid Cloud

The difference is structural. You own the hardware. Runpod orchestrates it.

FAQ

Questions a buyer actually asks.

Straight answers, including the edges. For anything deeper, you get a direct line to our engineers.

How does the security model work?

The host agent connects outbound over an encrypted overlay, with no inbound ports opened. Everything you onboard is scoped to your org and invisible to other customers. Bringing it to a security team? You get direct access to our engineers to walk the architecture.

What happens to my data?

Workloads run on your hardware and persistent data stays on local NVMe you own. The control plane only handles orchestration. If a workload overflows to Runpod cloud, that portion runs on Runpod infrastructure, so tell us early if you have strict data-residency requirements.

Is this lock-in?

No. The hardware stays yours, the host agent uninstalls, and your workloads are standard containers that run anywhere containers run. What you would give up by leaving is the orchestration layer, the same trade as any managed layer on hardware you own.

How many teams run this today?

We are early on purpose and keeping access deliberately limited, with white-glove onboarding and direct roadmap influence for the teams we work with. The orchestration fabric underneath already runs the Runpod cloud, so the foundation is mature even though this product is new.

Why pay to manage compute I already own?

The fee buys utilization and developer time on hardware that is already paid for. Idle owned compute is the most expensive line you have, and the orchestration and developer experience are what most teams are missing. The same layer is why teams pay a premium for Runpod's cloud.

How do pricing and procurement work?

Postpaid and enterprise-friendly: an annual management fee that scales with the compute under management, and professional services for onboarding. Pricing is set with you, and early customers get white-glove onboarding. Talk to us and we will shape it to your estate.

Can I run this on cloud instances I already have?

Yes. Any VM you control works, in a cloud account or on-prem, and we have demonstrated it on AWS and GCP. The host agent install is the same everywhere, and Runpod does not provision or touch your cloud accounts. You bring the hardware, we orchestrate it.

Pricing

Postpaid, and built for enterprise procurement.

You pay for the orchestration layer and the developer time it gives back, on hardware that is already yours. Idle compute is the expensive part.

Priced to your estate. We will walk the model with you.

Working with us

White-glove from day one, with our engineers in the room.

We are onboarding a small number of teams with hands-on support and real influence on where Runpod Hybrid Cloud goes next.

White-glove onboarding

We bring your hardware in and integrate with your environment, hands-on.

Direct line to engineering

Talk to the people who build it, fast.

Influence the roadmap

Early customers shape what we build and what comes next.

Migration support

We help move workloads onto your compute without a replatform.

On the roadmap

Where Runpod Hybrid Cloud is going.

We ship with design partners today. These are planned, not yet available, and shaped by the teams we work with now.

Serverless on your inventory

Serverless endpoints that target your hardware first, then burst.

Unified governance

Roles, audit trails, and budgets across owned and cloud capacity.

Cost attribution

Spend and usage attributed across your hardware and the cloud.

Fair-share allocation

Deploy tags and fair scheduling for shared clusters.

Runpod Hybrid Cloud

Put the compute you already own to work.

One control plane for the AI compute you own and rent. Your hardware first, Runpod cloud for overflow, the same developer experience across both.

Runpod, the AI Developer Cloud.

Request a demo