Brendan McKeagJune 11, 2025When to Choose SGLang Over vLLM: Multi-Turn Conversations and KV Cache ReusevLLM is fast, but SGLang might be faster for multi-turn conversations. This post breaks down the trade-offs between SGLang and vLLM, focusing on KV cache.AI InfrastructureAll
Brendan McKeagJune 6, 2025How to Deploy VACE on RunpodLearn how to deploy the VACE video-to-text model on Runpod, including setup, requirements, and usage tips for fast, scalable inference.AI WorkloadsAll
Brendan McKeagMay 31, 2025The 'Minor Upgrade' That’s Anything But: DeepSeek R1 0528 Deep DiveDeepSeek R1 just got a stealthy update, and it's performing better than ever. This post breaks down what changed in the 0528 release, how it impacts.AI WorkloadsAll
Brendan McKeagMay 23, 2025How to Connect Cursor to LLM Pods on Runpod for Seamless AI DevUse Cursor as your AI-native IDE? Here’s how to connect it directly to LLM pods on Runpod, enabling real-time GPU-powered development with minimal setup.AI WorkloadsAll
Brendan McKeagMay 16, 2025Automated Image Captioning with Gemma 3 on Runpod ServerlessLearn how to deploy a lightweight Gemma 3 model to generate image captions using Runpod Serverless. This walkthrough includes setup, deployment, and.AI WorkloadsAll
Brendan McKeagApril 30, 2025Qwen3 Released: How Does It Stack Up?Alibaba's Qwen3 is here, with major performance improvements and a full range of models from 0.5B to 72B parameters. This post breaks down what's new, how.All
Brendan McKeagApril 22, 2025Runpod Global Networking Expands to 14 More Data CentersRunpod’s global networking feature is now available in 14 new data centers, improving latency and accessibility across North America, Europe, and Asia.AI InfrastructureAll