
Which GPU should you actually use for embedding workloads?
Selecting the right serving engine for your embedding model can dramatically outperform hardware upgrades, yielding up to an 11x throughput increase on the same GPU.
Blog
Keep the platform your developers already use. Add reserved capacity, company-wide controls, compliance, invoicing, and hands-on support.

More than 1 million developers use Runpod because they can get compute in seconds, without taking a sales call. That isn’t changing.
Runpod Enterprise is for what comes next: when the same workload needs reserved capacity, a security review, company-owned administration, post-paid billing, and a contract the business can stand behind.
You shouldn’t have to choose between the cloud your developers already use and the cloud your company can approve. With Runpod Enterprise, teams can start self-serve, prove a workload in production, then formalize it under an enterprise agreement. The platform stays the same. The capacity, controls, support, and commercial terms grow with you.
What’s new this quarter:
Today’s announcement builds on Runpod’s foundation of production-ready infrastructure, including reserved and dedicated capacity across 30+ GPU configurations and 32 regions, with publicly reported region-level uptime. Customers also have contractual SLAs, named technical account managers and transparent incident communication.
Runpod Enterprise isn’t a separate enterprise cloud. It’s the same Runpod platform, backed by the commercial, operational, and governance model companies need when AI moves from experiments into core systems.
On-demand compute is the right default for experiments. Production workloads need more predictability, especially when a launch date, training milestone, or customer commitment depends on them.
Runpod Enterprise agreements give teams reserved and dedicated capacity for a committed term, planned against the regions and hardware their workloads need.
Capacity is not always simple. When a first-choice configuration is constrained, enterprise customers have a team that understands their requirements and can help find a workable path.
Reserved capacity is a supply commitment, not a promise that nothing will ever fail. No one can guarantee that. GPUs fail. Regions have bad days. Networks degrade.
Runpod Enterprise puts an operating model around that reality: contractual SLAs where they apply, support commitments with clear response expectations, and technical account management for onboarding, capacity planning, migration, and growth.
The contract matters because it makes expectations explicit. The relationship matters because production issues need context. Enterprise customers should not have to re-explain the workload, the capacity plan, and the deployment pattern every time they need help.
Ambitious AI teams rarely stay in one product. They may train on Clusters, iterate on Pods, and deliver inference on Serverless. Without an enterprise agreement, each motion can become a separate conversation about commitment, pricing, billing, access, and support.
Runpod Enterprise brings those workloads under one agreement, with usage metered and invoiced monthly at the organization level. One partner, one bill, one set of terms across development, training, and inference.
For teams already building on Runpod, the path from prototype to contract does not require a new platform. The same API and console stay in place; the agreement and controls catch up to how serious the workload has become.
Enterprise adoption depends on more than developer approval. Security and legal need documentation they can review, and finance needs terms they can process.
Runpod Enterprise supports SOC 2 Type II, HIPAA with BAAs where needed, GDPR with DPAs where needed, and ISO/IEC 27001 certification. Security review becomes part of the enterprise motion instead of a blocker that appears late in the deal.
The product proof sits underneath that motion: Organizations, SSO/SAML, access controls, billing visibility, resource tagging, centralized credentials, and audit trails. These are the controls that let the platform your developers already use become a company-owned environment.
Every enterprise agreement includes named technical account management.
That matters for production AI workloads. A support tier is not the same as an engineer who understands the workload, the capacity plan, and the deployment pattern. A shared inbox is not the same as someone who knows what changed since the last implementation call.
Runpod keeps technical support close to the work: onboarding, capacity planning, SSO setup, deployment design, migration, and scaling. When customer-specific work points to a broader product need, the goal is to bring that learning back into the platform rather than leave the customer with a one-off implementation.
Organizations are hierarchical, company-owned accounts. The organization owns its Pods, endpoints, Clusters, and volumes, rather than the individual who created them. When an engineer leaves, their work stays with the company.
SSO/SAML is part of that operating model. Enterprise SSO setup is handled with Runpod support, with ad hoc invites still available for contract users. Domain restrictions are coming next, and admin self-service for SSO is a longer-term roadmap item.
Every organization is private by default. Access requires two things:
Users may gain resource access by creating it, receiving a direct share, belonging to a group with access, serving as an organization administrator, or holding an organization-wide capability through their role.
Groups and policies let administrators tag resources with a group name and grant access based on membership. Permissions can cover actions such as injecting SSH keys into Pods, accessing storage through S3, or deploying and deleting inference endpoints.
Role-based access control provides standard roles and granular permissions across major access dimensions. Custom roles are on the roadmap. Today, administrators can assign permissions that are both specific and easy to explain.
Runpod also guards against privilege escalation. Users can grant only roles equal to or less privileged than their own, with checks applied when invitations are sent or accepted and when roles change.
Runpod Enterprise gives admins more visibility into spend, access, and account activity.
Cost Center and Billing Explorer provide organization-wide spend visibility, member-level drill-down, and cost-center labels for chargeback. Cost-center tags can appear on billable resources, and invoices can be grouped by cost center. That makes questions like "What did the Frankfurt inference workload cost last month?" answerable without rebuilding the bill in a spreadsheet.
Resource tagging is organization-wide, with system tags for region and data center applied automatically. Tags work in console resource views and API requests, so they become part of how teams organize infrastructure rather than cleanup work after the fact.
Credentials are now easier to manage at the organization level. SSH keys are individually addable and removable, and keys can be injected into group resources, so one person's key rotation does not break everyone else's access. Organization-wide shared secrets and container registry authentication come next, extending the same model to more credentials over time.
Audit Trails give each user a record of their own activity and give administrators an all-users view across the organization. Runpod is continuing to expand event coverage and API access, but the launch value is already useful: admins can see more of what changed, who changed it, and where to look when something needs review.
If you’re already using Runpod, you can initiate an enterprise conversation directly from the console.
Blog Posts