AI HUM TCH SU Sumit Vedpathak · Towards AI Your LLM Server Is Wasting 80% of Its GPU Memory — Here’s How vLLM Fixes That PagedAttention borrowed a 40-year-old idea from operating systems. The result: 24x higher inference throughput, same hardware.
AI MDA TCH SU Sumit Vedpathak · Coinmonks Your IDE Has a Secret Brain — GitHub Copilot in Visual Studio, Fully Dissected You press Tab. Copilot suggests. But do you know what’s actually firing under the hood? It’s four processes, a compiler bridge, and an agent
AI TCH HUM SU Sumit Vedpathak · Towards AI RAG from Scratch [Part 2]: Loading — The Step Everyone Skips and Everyone Regrets Series 2 of 5: The unglamorous first step that quietly decides whether your entire RAG pipeline succeeds or silently fails.