MDA AI EM Ema Ilic Four days ago, a study was published that really got me thinking. Dynamic Memory Compression as a solution to low inference speed in pre-trained LLMs.