SOC AI FA Falak Shair I Wrote My First GPU Kernel — Here’s What Changed 1.49x faster softmax. But only 1.074x faster forward pass. Amdahl’s Law explains the gap — and understanding it matters more than the…
HUM AI IR Irfan Mansuri Why Batching Hurts Reasoning (and Why Systems Still Do It) Batching is one of the most effective optimizations in deep learning. It increases arithmetic intensity, hides latency, and maximizes GPU…