<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Vllm on Kashish Blog</title><link>https://www.kashishverma.com/tags/vllm/</link><description>Recent content in Vllm on Kashish Blog</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Wed, 26 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.kashishverma.com/tags/vllm/index.xml" rel="self" type="application/rss+xml"/><item><title>LLM Inference at Scale: vLLM and llm-d</title><link>https://www.kashishverma.com/posts/til-2/</link><pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate><guid>https://www.kashishverma.com/posts/til-2/</guid><description>&lt;p&gt;Running AI infrastructure breaks a lot of the traditional ways we&amp;rsquo;re used to dealing with systems. In this blog, I write about how you can serve AI models efficiently for inference on Kubernetes on hardware like GPUs (which are super expensive!!!), and how vLLM can help you manage memory efficiently while llm-d can help load balance the system smartly to your needs.&lt;/p&gt;</description></item></channel></rss>