SPT AI AL Allen Kuo (kwyshell) Finishing What We Started: Gemma 4 NVFP4 on vLLM, Desktop Blackwell, WSL2 A follow-up to Gemma 4 on vLLM vs Ollama: Benchmarks on a 96 GB Blackwell GPU.
AI SPT SE Sebastien Running Mistral Small 4 (119B, NVFP4) Locally on a DGX Spark A practical guide to serving a 119B parameter model on NVIDIA’s compact Blackwell workstation — including the quirks of the consumer GB10…
AI TCH TH Thomas P. Braun · Avarok We Unlocked NVFP4 on DGX Spark — and It’s 20% Faster Than AWQ NVIDIA’s own software stack couldn’t do it. We did. Avarok’s open-source vLLM image is the first to make NVFP4 outperform AWQ on GB10.
AI SCI SPT SA Sarankannan Bleeding Edge or Bleeding Out? The Quest for vLLM on NVIDIA Blackwell Why my 57GB BF16 experiment failed, and what it reveals about the gap between hardware delivery and software reality.
AI MD Md Monsur ali · Data Science Collective NVIDIA NVFP4: LLM 4-Bit AI Training Breakthrough Explained NVIDIA’s NVFP4 format enables efficient 4-bit LLM training with 12B parameters on 10T tokens, achieving a 3x speedup with no loss in…
AI HUM DI Dinmay kumar Brahma Pretraining Large Language Models with NVFP4: An In-Depth, Accessible Guide Large Language Models (LLMs) power cutting-edge AI — from chatbots like ChatGPT to generative code tools and translation engines. Training…