News icon

Kimi K3 is now available on Runpod

Make the model yours

Customizability is the most underrated idea in AI right now. Runpod CEO Zhen Lu on why a model tuned on your data beats a bigger one on the job you actually have.

Make the model yours

Civitai's community trains 868,000 LoRAs a month on Runpod.

I used that number in my last post and I haven't stopped thinking about it. A LoRA is a small adapter that teaches a base model something specific about your subject or your style, trained in hours on modest hardware. At that volume, customization is mass behavior. It stopped being a research technique and became something hundreds of thousands of people just do.

That post was about control: when you own the weights, nobody can deprecate, reprice, or reroute the model you shipped. This one is about the second thing teams keep weighing, the ability to make a model their own. Inside Runpod we shorthand the full set as the 3 Cs: control, customizability, and cost.

Customizability is the most underrated idea in AI right now.

A tuned model wins on the task you actually have

Most teams still assume the best model for any job is the biggest general one. That holds on a benchmark of everything, and it stops holding surprisingly often on a single task you understand well.

Predibase ran the cleanest version of this experiment I've seen, a study called LoRA Land: 310 fine-tuned models in the 7B class, evaluated across 31 tasks. The fine-tuned small models beat GPT-4 on 85% of those tasks, by ten points on average.

GPT-4 is ancient history now, but the mechanism hasn't aged. A study this June found fine-tuned models under a billion parameters beating GPT-5.4 zero-shot on document extraction. Predibase sells fine-tuning tools, so discount accordingly. The result still matches what we see on our own platform every day: a model that has trained on your data carries an advantage that general capability doesn't replace.

And the advantage compounds. A general model treats your domain as one slice of everything it knows. A tuned model treats your domain as the whole job, and every round of your data makes it better at exactly the thing you ship.

The starting point got good

Fine-tuning open weights used to mean accepting a much weaker base and earning the difference back with your data. That tradeoff has collapsed.

Stanford's AI Index measured the gap between the top closed model and the top open one on Chatbot Arena, a public leaderboard where users vote on model answers blind: 8% in January 2024, 1.7% thirteen months later. Epoch AI's read this year puts the best open models about four months behind the closed frontier.

That number is already under pressure. Moonshot's Kimi K3, released in late July as open weights, trades blows with the top closed models on public leaderboards, the closest open models have been to the frontier since DeepSeek R1.

Four months matters if you're selling a general-purpose assistant. For a scoped task where your fine-tune does the differentiating, it's noise.

Deep customization now lives on open weights

The market is reorganizing around this. Two weeks ago Thinking Machines shipped Inkling, its first open-weight model, built as raw material for fine-tuning. The launch coverage described a Bridgewater project where a model tuned on the firm's own research beat frontier models on its internal work, at a fraction of the cost to run.

OpenAI, meanwhile, is winding down self-serve fine-tuning: closed to new organizations in May, narrowed again in July, no new training jobs after January 2027. I don't know the reasoning behind it. The practical fact for builders: if making a model your own is part of your plan, open weights you can hold are where that work lives now.

Start with a frontier model, still

None of this changes the advice I gave last month. Most teams should start with frontier models.

The strongest case against fine-tuning is that the frontier keeps moving. Whatever you tune this quarter, a general model may match next quarter, and prompting plus retrieval will carry you a long way without a single training run. That case is right more often than my side of the industry likes to admit. Fine-tuning without a scoped problem and usable data is how a team burns a quarter producing a model that performs worse than the prompt they started with.

The graduation moment is the same one I described last month: a scoped problem, data you can use or a plan to create it, or security requirements that rule a closed model out. Reach that point and the calculus flips. The tuned model wins on the task today, and it keeps compounding, because every improvement you make is yours to keep.

What I don't know

I don't know whether the gap keeps closing. Epoch's own data showed it widening this spring, from about three and a half months to four, and then K3 swung it the other way. The frontier labs may pull away again, and that would change the base you start from. I don't think it changes the argument. The case for customization rests on your task and your data, and those don't belong to any lab.

The next post covers the last of the three, cost. If any of this doesn't hold up against what you're seeing in your own work, I want to hear it.

Related articles

View All

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.

Star field background