AI HUM SH Shirley Li · AI Advances DeepSeek-V3 Explained 3: Auxiliary-Loss-Free Load Balancing How DeepSeek breaks the hidden bottleneck in MoE
AI HUM SH Shirley Li · Data Science Collective DeepSeek-R1: Advancing LLM Reasoning with Reinforcement Learning This is the seventh article in our DeepSeek series [1], where we will break down how DeepSeek-R1 is trained by exploring large-scale…
SPT HUM AI SH Shirley Li MatFormer: Train Once and Deploy Many The secret weapon of Gemma 3n to let one model fit all.
AI HUM SH Shirley Li · Data Science Collective DeepSeek Explained 6: All you need to know about Reinforcement Learning in LLM training This is the sixth article in our DeepSeek series, where we will dive deeper into one of the key innovations in training strategies of…
AI HUM SH Shirley Li · Data Science Collective DeepSeek Explained 4: Multi-Token Prediction How DeepSeek achieves better balance between efficiency and quality in text generation
AI HUM SH Shirley Li DeepSeek Explained 8: Post-Training of DeepSeek-V3 This is the last article of our DeepSeek series ([1], [2]), finally we will cover the post-training techniques in DeepSeek-V3 [2].
AI HUM SH Shirley Li · Data Science Collective DeepSeek Explained 5: DeepSeek-V3-Base Innovations in pre-training strategies of DeepSeek-V3.
AI SH Shirley Li Transforming the Transformer (Part 1): The Rise and Fall of Efficient Attention A deep dive into the rise and fall of X-formers, and why modern LLMs abandoned most of them.
AI ART HUM SH Shirley Li Paper Explained 3: E5 How a simple architecture is transformed into a SOTA embedding model
AI HUM SH Shirley Li Paper Explained 4: NV-Embed How Nvidia turns decoders into efficient embedding models