FEATUREDQbitAI (量子位) · WeChat· rssZH04:04 · 05·16
→Zhejiang University and Microsoft use 3,000 text prompts to improve video 3D consistency with World-R1
Zhejiang University and Microsoft introduced World-R1, training Wan 2.1 with about 3,000 text-only prompts, Flow-GRPO, and a four-part reward; the 1.3B version improves PSNR over the baseline by 10.23 dB.
#Multimodal#Vision#Alignment#Zhejiang University
why featured
Featured · importance 78 · hook + knowledge + resonance
editor take
World-R1’s trick is not “teaching 3D”; it pressures Wan 2.1 with 3,000 prompts and rewards. Useful, but don’t call it a world model.
sharp
World-R1 is a reward-engineering win, not proof that video models suddenly understand physics. On Wan 2.1, it changes no architecture and uses no 3D data; it trains with about 3,000 Gemini-written prompts, Flow-GRPO, and a four-part reward. The numbers are real: the 1.3B model gains 10.23 dB PSNR over baseline, and LPIPS drops from 0.467 to 0.201.
I don’t buy the “3D knowledge was asleep” framing. The reward stack uses Depth Anything 3, 3D Gaussian Splatting, Qwen3-VL, and HPSv3, so a lot of visual prior sits in the judge. The clever bit is encoding camera trajectory into initial diffusion noise instead of adding a control net, while beating ReCamMaster and DAS on aesthetic scores. The unresolved risk: reward overfitting to reconstruction metrics; cross-model replication is not shown here.
HKR breakdown
hook ✓knowledge ✓resonance ✓