SPT MDA AI MK M K Pavan Kumar Compressing Gemma 4 12B with LLM Compressor and Serving It on vLLM This is the full journey of taking Gemma 4 12B from a 23.9GB BF16 model to a 15GB FP8 one that serves on a single GPU — start to finish…
AI MDA JI JINIMINI Why LLM Compression Matters Today In recent years, Large Language Models (LLMs) have grown rapidly in size and capability. Models with billions (or even trillions) of…
MDA AI VI Vicens Gaitan Structural Pruning in Mixture of Experts (MoE) Models: Applied Analysis to GPT-OSS-120b
MDA AI YO YouShin kim [Google Research] — TurboQuant: 초고압축 기술을 통한 AI 효율성의 재정의 1. 배경 및 개요 (Background and Overview)
TCH MDA AI KA Kawaldeep Singh On-Device LLMs in 2026: How Teams Ship Fast, Private, and Tiny Large-Model Features In 2026 you no longer need a multi-GPU server to deliver useful LLM experiences. The new game is about squeezing reliable language features…
AI MDA AN Aniket Sanyal · Towards AI Demystifying DPKD: How Preference Knowledge Distillation Boosts Small AI Models 🚀 Introduction: Big Brains vs Small Brains in AI 🧠
MDA AI SPT SI Siemghirmai LLMs At the Edge: Compressed Language Models Large Language Models (LLMs) are shifting from data centers to smaller devices, making advanced language processing more accessible on…