News icon

Kimi K3 is now available on Runpod

How we think about pricing at Runpod

Our pricing philosophy in one line: we move prices to keep GPUs available.

How we think about pricing at Runpod

If you run workloads on Runpod, you’ve probably noticed GPU prices move month to month. A card you rented in the spring costs a little less today. Another costs a little more. Prices move because supply and demand move, and because availability is what matters most to us in the end. A low price doesn’t help much if the GPU is never available; a high price on an underused GPU leaves useful capacity sitting idle.

Static prices sound like a promise. In a supply-constrained market, it’s a promise no one can keep.

Why prices move at all

GPUs are among the most supply-constrained hardware in computing, and demand for them arrives in spikes. When a new open model drops, thousands of teams can want the same card in the same week. New capacity takes longer to add, often months.

When demand outruns available GPUs, capacity gets rationed whether we call it that or not. That rationing usually shows up in worse ways: capacity errors, queues, waitlists, sales gatekeeping, multi-year contracts. Those work fine if you have a legal team and a seven-figure budget. They work badly if you’re one developer with a credit card and something to ship.

Runpod started by building for that developer: deploy in under 30 seconds, without a sales call or contract. That self-serve access isn't going anywhere, even as we grow.

There's a second, more important reason. On-demand hourly GPUs compete for the same silicon as multi-year committed contracts. If you hold on-demand prices below what a committed buyer will pay, that supply goes to the committed buyer. The on-demand pool thins out. Pricing on-demand at something close to its market value is what keeps an on-demand pool there at all.

In short? Our prices move so that access doesn’t become a queue.

How adjustments work in practice

We review pricing monthly, per GPU type. Utilization across the fleet is the main input, alongside shifts in what it costs us to secure that capacity, but the emphasis is on demand: what you and every other developer on the platform are asking for in a given month. If a GPU is consistently underused, we lower the price. If it’s under sustained pressure, we raise it. The price adjustments are intentionally small. They should reflect demand without making your bill feel unpredictable.

One thing keeps this from turning into whiplash: we aim to change prices infrequently rather than continuously. Less often could make each individual change seem bigger, but we think it’s a better move to absorb a few legible changes a year than dozens of small ones you can't plan around.

Just as important is what doesn’t change:

  • You pay for compute that's actually running. Initialization and errors don't count against you. Provisioning time does, since it scales with the size of your image, not with demand.
  • Serverless scales to zero. Idle costs nothing.
  • There are no egress fees on S3-compatible storage. Your data is never a hostage to your bill.
  • Every price is public at runpod.io/pricing. No call, no custom quote, no NDA.

Is this sustainable?

It’s fair to ask. Cloud pricing has trained us to look for the catch: free credits that disappear after migration, intro rates that turn into enterprise rates, discounts that only exist once you’ve committed. All of those work the same way: get you in cheap, earn it back later.

That isn’t the model here. Our prices reflect the unit economics of running AI workloads on a platform built specifically for that job. We’re not carrying the overhead of a general-purpose cloud, and we pass that difference through, which is why compute costs on Runpod can run up to 90% lower than traditional cloud providers.

The prices you see aren't an intro rate we'll need to claw back later. If you outgrow what self-serve covers, scaling further means a straightforward conversation about a plan that fits your usage, not us padding your bill to make up for that early rate.

What this means for your bill

If your workloads can run on more than one GPU type, monthly adjustments can work in your favor. Markdowns usually show up where there is available capacity, and for a lot of jobs that card ends up being the best price-performance option on the platform. If you need the most in-demand card during the week everyone else wants it too, you’ll pay closer to what it’s actually worth once the next monthly review hits. It’s not real-time surge pricing. What you’re quoted today holds until then. The tradeoff is that you can get it now, without a contract.

Where this goes next

If something about your bill surprises you, that’s on us. Send it to [email protected] with the pod or endpoint ID and roughly when the charge landed, and we'll walk through it with you. Nobody should feel tricked by their infrastructure, least of all the part they pay for.

Related articles

View All
When (and why) to upgrade your Python version

When (and why) to upgrade your Python version

Python versions carry a security clock and a performance upgrade you're leaving on the table — here's what changes when you move off an old interpreter, and when it's fine to wait.

All

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.

Star field background