News icon

Kimi K3 is now available on Runpod

What 1.2 million developers taught me about AI this year

The State of AI Compute shows what teams actually deploy: agents are 8% of users but drive 24% of revenue, a new model version can take over its family's workloads in about a month, and lower precision is becoming a production strategy for large models.

What 1.2 million developers taught me about AI this year

Most of what gets written about AI adoption comes from surveys. Someone asks a thousand executives what they plan to do, and we get a chart of intentions.

My job is different. At Runpod, developers pick the model, the precision, the GPU and the runtime, one workload at a time, and the platform records those choices. Across 1.2 million accounts in 186 countries, that record is the largest dataset I've seen on what teams actually run when compute is scarce.

We just published the second edition of our State of AI Compute report. It's full of charts. But the numbers that stay with me are the ones that surprised our own team. Here are seven of them.

1. Agents are 8% of our users and 24% of our revenue

This is the one I keep coming back to.

Across three consecutive monthly snapshots, the share of Runpod revenue tied to resources created by agents, not humans, went from 10% to 12% to 24%. Over the same period, agents grew from 2.4% to 4.2% to 8.4% of users. Per user, agent-created resources now bring in about 2.9x the platform average.

A caveat, because I'm a data person: the agent isn't the customer. A human or a company still owns the account and pays the bill. What changed is who, or what, initiated the deployment. Increasingly, it's software.

The part I find more telling is how long these workloads last. Agent Pods are 3.7x more likely to run longer than a month than the platform baseline (7.7% vs. 2.1%), and much less likely to disappear within an hour. Longer runtime doesn't prove that a workload is in production. But things that run for weeks usually aren't demos.

2. Most AI video workloads don't generate any footage

Ask people what AI video is, and they'll probably describe text-to-video. Our data says otherwise.

Among video Pods, 62.9% only work on existing footage. They never create a new frame from scratch. Overall, 85.2% do some kind of post-production, and upscaling (77%) is the most common workflow by a wide margin. Only 11.2% generate video without post-processing it.

Image workloads show the reverse pattern: generation shows up in 82.2% of image Pods. Audio is mostly transcription, with speech-to-text appearing in 43.6% of audio Pods.

The lesson is that "generative AI" isn't one workload. A video upscaler and an image generator have different memory, storage and latency requirements. If you plan infrastructure around the label instead of the workflow, you'll buy the wrong thing.

3. Open and frontier models are built into the same systems

Most coverage of open versus frontier models reads like a scoreboard. Our data shows something else: among Pods that call a frontier API, 80% also run open weights in the same environment.

The split often follows specialization. Frontier APIs handle the general-purpose parts of an application. The decisions specific to a business go to open models fine-tuned on that business's own data. On Runpod, small specialist models already handle jobs like redaction, product tagging and validation, and they're 2.7x more likely to be fine-tuned. Gartner projects task-specific models to be used three times as much as general-purpose LLMs by 2027.

A general model is built to be good at everything. A fine-tuned one is built to be right about your problem.

4. Teams running their own weights upgrade within weeks

With an API, upgrades happen for you: the provider ships a new model and you get it. Running your own weights means doing that work yourself, which is why many teams assume self-hosted deployments get stuck on older versions. Our data shows the opposite: when a better open model ships, teams move to it within weeks.

Kimi is the clearest example. In July, K2.7 was the most common version among self-hosted Kimi Pods, at 45%. By August, the newly released K3 held 60%. Qwen's share of text endpoints jumped from 61% to 69% in the month its Qwen3.5 weights were released, and Gemma went from 5.5% to 18.2% after the release of Gemma 4.

Add up enough of those switches and the whole field changes. Llama, the default open text model two years ago, fell from 23% of text endpoints in November 2025 to 7.6%. Qwen now runs on 74.2%.

5. Teams are running models at lower precision

Tight supply didn't stall teams. It pushed them to make compute go further. In a year when high-bandwidth memory was booked into 2027 and flagship GPU prices went up, teams started moving away from full precision.

Among Qwen3 endpoints on Runpod, the share running 16-bit builds fell from 95.9% to 85.8% over the period we studied. Over the same period, 8-bit builds rose from 2.0% to 10.9% and 4-bit from 3.1% to 8.8%.

The shift is strongest where memory matters most. 40.7% of Pods serving models above 70B parameters use a quantized build, compared with 11.6% of Pods serving models under 8B.

6. AI is becoming standard software, everywhere

AI is no longer an experimental technology. It's production infrastructure, running in every kind of business and nearly every country.

AI companies are still the core of our business: 30.2% of users and 61.6% of revenue. But they're only 25.9% of new business signups, and eight of our ten fastest-growing industries aren't AI-specific. The top three are government, fintech and manufacturing. AI isn't their core business, but they're building with it anyway.

It's spreading geographically too. We now have users in 186 countries. The U.S. accounts for 18.6% of users and 38.8% of revenue. Asia is our fastest-growing continent, and Japan is our third-fastest-growing country. 

In our data, new markets tend to add users before they add spend. If you're looking for early signals on where demand is heading, I'd keep my eyes on Asia.

7. We graded our own forecast

Last year we predicted that H100 and H200 supply would roughly double by mid-2026 and B200 would nearly quadruple. Here's how those forecasts held up:

  • H100 SXM: Supply roughly doubled. Close to the forecast.
  • H200: Supply grew about 1.6x, short of our forecast.  .
  • Blackwell: B200 supply tripled instead of quadrupling. We didn't forecast B300, which came online during the year. Counting both, our Blackwell fleet grew about 6x, beating the 4x we forecast. Our miss was not seeing how fast the next generation would arrive.

I include this because a data team that only reports its wins isn't worth trusting. Forecasts are hypotheses. Checking them in public is how we get better at making them.

Why this matters

As AI moves into production infrastructure, teams are making decisions based on performance, efficiency and control over their own workflows. That means optimization is part of the job now. Precision, model size, hardware class and serving pattern decide cost, latency and quality, and the right answer changes whenever supply or models change.

Most published research about AI relies on surveys and forecasts. We see deployment decisions directly: which models, precision, and GPUs teams choose across 1.2 million accounts in 186 countries. Agents creating their own infrastructure, AI spreading beyond the AI industry, new markets adding users before spend: these shifts show up in our data as they happen.

We don't have to predict where AI is going. We watch it as it evolves, and we build for what's next. The next big shift in AI is already running on Runpod.

‍

The full report, with methodology and denominators for every number, is here. If there's a question you think our data could answer, send it my way. Better yet, send the one you think it can't.

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.

Star field background