News icon

We raised a Series A! Read a post from our CEO, Zhen Lu: 1M devs and the cloud we're building next.

Custom models are a control decision

Every rented model comes with someone else's roadmap. Runpod's CEO on why the smartest teams are starting to build their own.

Custom models are a control decision

Teams building serious AI products eventually want three things from their models: predictable costs, control over how they behave, and the ability to make the model their own. Our team calls them the 3 Cs. Recent events have made control hard to ignore.

What just happened

On June 9, 2026, Anthropic shipped Fable 5, the most capable model it has ever made generally available. Within a day, developers found capability limits for certain workloads that had only been disclosed deep in the system card. They also found a downgrade tag in system logs that read TOO_DUMB_TO_NEED_FABLE. The new model class came with a 30-day data retention requirement and pricing of $10 per million input tokens and $50 per million output tokens, double the previous top model.

On June 10, Anthropic made the limits visible and apologized: “We made the wrong tradeoff.” They fixed the issue within a day and deserve credit for responding quickly. But developers still had no control over what they were being served.

The same thing has been happening at OpenAI on a slower timeline. In February 2026, OpenAI retired GPT-4o from ChatGPT and ended API access to the endpoint that tracked it, with the final snapshot scheduled to shut down in October 2026. Teams that spent months tuning latency-sensitive production pipelines around the model were given a deadline to migrate. Moving a tuned production system is real work, and someone else controlled the timeline.

OpenAI has also been closing access to self-serve fine-tuning in stages. New organizations were cut off on May 7. Access narrowed again on July 2. On January 6, 2027, even active customers will lose the ability to create new training jobs.

I won’t guess at the reasoning behind these decisions. Frontier labs have to balance safety, capacity, and economics across millions of users, and they generally do that job impressively well. But their constraints are different from yours. When your product depends on a rented model, the provider controls the price, version, behavior, data policy, and retirement date. You need a plan for when one of those things changes.

Start with frontier models

Most teams should start with frontier models. I use them every day. They give you the most capability for the least engineering effort, and they are usually the fastest way to find out whether an idea works.

If you are still figuring out the problem, renting the best available intelligence is usually the right decision.

That starts to change once the use case becomes clear. When you understand the problem and have data around it, or a credible way to create that data, fine-tuning an open source model becomes much more attractive. The same applies from day one when security or data requirements rule out a closed model.

At that point, you can stop paying for a broad set of capabilities you may not need. You can build a model around your domain and decide where and how it runs.

What control means

Control means the model you ship is the model you tested, and you can keep running it for as long as you choose. You control the weights, version, deployment, and data path. It cannot be deprecated, repriced, or changed without your knowledge.

The weights are only part of it. You also need control over how the model is served, observed, stored, and governed. That requires real engineering.

A lot of the current conversation about AI infrastructure focuses on token prices and rate limits. Across the million-plus developers on our platform, many of the teams furthest ahead are already fine-tuning and serving models themselves. Civitai’s community alone trained 868,000 LoRAs in a single month, with each one adapting a base model to a particular style or subject. Developers already want to shape and own their models, and they are doing it at scale.

Prompt engineering is a phase

I think prompt engineering is a phase. The durable work is building, deploying, and maintaining AI systems.

The developers who treat models as components they can shape and operate are building products that will last. They will also be much better prepared the next time a lab changes a model, raises a price, limits access, or shuts something down.

I may be wrong about how quickly teams make this move. If frontier model prices keep falling and capabilities keep improving, renting will remain the right choice for longer than I expect. It may take more time before the added control is worth the engineering work.

If you think I have this wrong, tell me where.

Control is the first of the 3 Cs. Cost and customizability will get their own posts, and I plan to publish one every two to three weeks. Together, they explain a lot of why we built Runpod. Developers should be able to build and run AI systems that are actually theirs.

The most interesting systems on our platform haven’t been built yet.

If any of this resonates, come give us a try.

Related articles

View All

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.

Star field background