James SandyNovember 6, 2025How to Run Serverless AI and ML Workloads on RunpodLearn how to train, deploy, and scale AI/ML models using Runpod Serverless. This guide covers real-world examples, deployment best practices, and how.Product UpdatesAll
James SandyApril 21, 2025How to Fine-Tune LLMs with Axolotl on RunpodLearn how to fine-tune large language models using Axolotl on Runpod. This guide covers LoRA, 8-bit quantization, DeepSpeed, and GPU infrastructure setup.AI WorkloadsAll
James SandyApril 14, 2025Cost-Effective AI with Autoscaling on RunpodLearn how Runpod autoscaling helps teams cut costs and improve performance for both training and inference. Includes best practices and real-world.AI WorkloadsAll
James SandyMarch 18, 2025Deploying Multimodal Models on RunpodMultimodal models handle more than just text, they process images, audio, and more. This guide shows how to deploy and scale them using Runpod’s infrastructure.AI WorkloadsAll
James SandyNovember 22, 2024How Much Can a GPU Cloud Save You? A Cost Breakdown vs On-Prem ClustersWe crunched the numbers: deploying 4x A100s on Runpod's GPU cloud can save over $124,000 versus an on-prem cluster across 3 years. Learn why cloud beats.Cost OptimizationAll
James SandyNovember 12, 2024Quantization Methods Compared: Speed vs. Accuracy in Model DeploymentExplore the trade-offs between post-training, quantization-aware training, mixed precision, and dynamic quantization. Learn how each method impacts model.AI WorkloadsAll