AI SU Sujangyawali How DeepSeek exactly implemented Latent Attention | MLA + RoPE DeepSeek’s Multi-Head Latent Attention (MLA) is one of the most important innovations behind DeepSeek-V2 and DeepSeek-V3. Instead of…
AI ECO AG Agent Native Founder’s Open-Model Stack: GLM-4.7, Qwen3-VL, DeepSeek-V3.2, Kimi-K2, FLUX.2 If you’re building an AI product as a solo founder or a small team, you don’t need one “best” model.
AI SCI SU Supat Charoensappuech Möbius Correction to αTP: Why the Geometric Value 0.08537 Is Not the Final Physical Value By Supat Charoensappuech (with assistance from DeepSeek-V3.2 and ChatGPT-5.1 in normal mode, 30/11/2025)
HUM AI SPT CH Chandan Kumar · Level Up Coding Deploying Deepseek 3.2 Exp on Nvidia H200 — Learning lessons This is a hands-on log of getting DeepSeek-V3.2-Exp (MoE) running on a single H200 box with vLLM. It covers what worked, what didn’t, how…
AI HUM SH Shirley Li DeepSeek Explained 8: Post-Training of DeepSeek-V3 This is the last article of our DeepSeek series ([1], [2]), finally we will cover the post-training techniques in DeepSeek-V3 [2].
AI HUM SH Shirley Li · Data Science Collective DeepSeek Explained 5: DeepSeek-V3-Base Innovations in pre-training strategies of DeepSeek-V3.
AI HUM SH Shirley Li · AI Advances DeepSeek-V3 Explained 3: Auxiliary-Loss-Free Load Balancing How DeepSeek breaks the hidden bottleneck in MoE
SCI TCH AI MED MDA SOC DA dave ginsburg · AI.society AI Society for 1.9.25 — Your Personal AI Supercomputer Today: Nvidia Project Digits and Cosmos World Foundation Models, Seagate’s HAMR, Large Concept Models, DeepSeek-v3 dangers, AI and…