Kimi K3 is now available on Runpod
A 2.8T-parameter, 104B-activated MoE with hybrid linear attention, native vision, and a 1M-token context. This covers the architecture that matters for serving it, how it differs from K2, and how to run it on Runpod. Details are drawn from the Kimi K3 technical report and Moonshot's K3 blog.
Size GPUs for browser-use and computer-use agents on Runpod, covering DOM-based vs screenshot-based architectures, VLM sizing, and colocation.
Deploy the Letta (MemGPT) server on Runpod and point it at a self-hosted vLLM or TGI model, with Postgres persistence and embedding model setup.
Point the OpenAI Agents SDK at a self-hosted vLLM or TGI model on Runpod using the built-in LiteLLM adapter, split across CPU and GPU deployments.
Deploy F5-TTS voice cloning on Runpod as a callable API, covering licensing, GPU sizing, and wrapping the Gradio demo in FastAPI for production.
Fine-tune DeepSeek V3's 671B MoE architecture on Runpod Clusters with LoRA, covering hardware sizing, framework choice, and FP8 numerics.
Fine-tune GLM-5 on Runpod with Axolotl, covering MoE-specific settings, GPU sizing, and a representative QLoRA config for expert-routed models.
Size the right GPU for LoRA, QLoRA, or full fine-tuning on Runpod, with a sizing table by model size and method, plus when to use Clusters.
Deploy Aphrodite Engine on Runpod for EXL2, GGUF, and GPTQ checkpoints, with AGPL licensing notes and modern sampler configuration.
Run the LiteLLM proxy on Runpod and route it to a self-hosted vLLM or TGI endpoint alongside cloud providers, with spend tracking and fallback.
Deploy LocalAI on Runpod with the correct GPU image, model gallery installs, and verification steps to confirm GPU acceleration is active.
Deploy LMDeploy on Runpod with TurboMind: GPU sizing, KV cache tuning, tensor parallelism, and quantization for AWQ and MoE checkpoints.
What an MCP server is, how Streamable HTTP transport works, and how to build and host your own MCP server on Runpod.
Compare GPUs for AI video generation with Wan 2.2, HunyuanVideo, and LTX-2.3, including VRAM requirements and Runpod pricing fit by use case.
Deploy Hugging Face TGI on Runpod: GPU sizing, Docker setup, Network Volume caching, production flags, and an OpenAI-compatible endpoint.
Learn how AI inference and training differ in GPU requirements, VRAM sizing, and cost, and how to match each workload to the right hardware in production.
Learn how vLLM boosts LLM inference performance with PagedAttention and continuous batching. This guide covers KV cache optimization, GPU efficiency, and.
Learn how to run SGLang in production with structured generation, RadixAttention, and multi-step LLM pipelines. Boost throughput by reusing KV cache and.
LLaVA 1.7.1 combines vision and language in one powerful model. This guide shows you how to get it running on Runpod in minutes.
Accelerate machine learning deployment with automated MLOps pipelines on Runpod, streamline data validation, model training, testing, and scalable.
Shares hacks to optimize GPU hosting for high-performance AI, potentially speeding up model training by . Explains how Runpod's quick-launch GPU.
Avoid costly GPU mistakes. Learn how to size VRAM, choose the right cloud instance, and deploy AI workloads efficiently without wasting budget.
Deploy distributed AI training across global cloud regions with Runpod, optimize cost, performance, and compliance using spot instances, gradient.
Learn how to deploy vLLM with Docker on Runpod end-to-end: from GPU selection and pod configuration to Network Volume caching, server flag tuning, and a.
Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.