AI LJ L.J. Bridging the “Spatial Blindness” Gap: How VEGA-3D Leverages Video Generation for 3D Understanding Modern Multi-modal Large Language Models (MLLMs) are surprisingly adept at “chatting” with images, yet they often hit a wall when faced…