SOC AI FA Falak Shair I Wrote My First GPU Kernel — Here’s What Changed 1.49x faster softmax. But only 1.074x faster forward pass. Amdahl’s Law explains the gap — and understanding it matters more than the…
AI HUM SPT FA Falak Shair I Profiled LLM Inference From First Principles — Here’s What I Found “Why is this so slow?” My colleague was frustrated. We work at the same company, and he was building a RAG system with a local LLM running…
AI JA Jaideep Ray · Better ML How Big Is an LLM? Count the Facts It Remembers Frontier labs rarely disclose parameter counts. As a result, practitioners rely on indirect proxies such as API latency, pricing, and…
HUM AI VI Vincent Chen ML Systems — Tensorflow & PyTorch This week we move on to the most famous ML frameworks— Tensorflow’s original paper and PyTorch 2!
AI SPT HUM VI Vincent Chen ML Systems — Introduction This quarter, I am enrolled in a Machine Learning Systems course. The following series of articles will serve as my summaries and…
HUM AI AL Aleksa Gordić ELI5: Flash Attention Step by step explanation of how one of the most important MLSys breakthroughs work — in gory detail.