DeepMind Turns 7,500 Synthetic Clips Into a 14B-Param Multi-Task Vision Model
The model, built on Alibaba’s Wan2.1 architecture, learns depth, normals, segmentation and 3D pose from a tiny Blender-rendered dataset, matching or beating specialist systems while needing only a fraction of the training data.