AI SCI AN Anuva Sharma The Journey to Multi-Head Latent Attention Why DeepSeek-V2 invented a strange-looking attention block — and how a tiny algebraic trick made the KV cache 57× smaller without dropping…