Marut PandyaSeptember 10, 2024Boost vLLM Performance on Runpod with GuideLLMLearn how to use GuideLLM to simulate real-world inference loads, fine-tune performance, and optimize cost for vLLM deployments on Runpod.AI WorkloadsAll
Marut PandyaJuly 1, 2024AMD MI300X vs. Nvidia H100 SXM: Performance Comparison on Mixtral 8x7B InferenceRunpod benchmarks AMD's MI300X against Nvidia's H100 SXM using Mistral's Mixtral 8x7B model. The results highlight performance and cost trade-offs across.All
Marut PandyaJuly 1, 2024AMD MI300X vs. NVIDIA H100: Mixtral 8x7B Inference BenchmarkWe benchmarked AMD’s MI300X against NVIDIA’s H100 on Mixtral 8x7B. Discover which GPU delivers faster inference and better performance-per-dollar.All