NVIDIA Physis-Lang: Language-Driven Physics for Realistic Video Generation
NVIDIA, MIT, and Oxford researchers propose Physis-Lang, a framework that embeds detailed physical reasoning into language captions for video generation, using a self-evolving benchmark and language-guided data retrieval to achieve state-of-the-art results on physics benchmarks, surpassing Veo 3.1 on three of four tests.
Physis-Lang: Language as Physics Representation for Video Generation
Researchers from NVIDIA, MIT, and Oxford University introduce Physis-Lang, a framework that embeds detailed physical reasoning — causes, laws, and outcomes — into natural language captions to improve the physical realism of generated videos. The approach operates across three stages: data selection, model training, and inference.
Physics-Aware Captioning and PhysCapBench
Standard video captions describe scenes and actions but omit the physical processes linking cause and effect. Physis-Lang enriches captions with continuous state changes. For example, instead of "butter melts," a Physis-Lang caption reads: "Continuous heating gradually melts the butter; the solid portion collapses under gravity, shrinking in volume, while the melted liquid forms a shallow pool that expands outward." This forces the model to represent temperature-driven phase transitions, gravity-driven deformation, and concurrent solid–liquid dynamics.
To evaluate caption quality, the team built PhysCapBench , comprising 246 physics videos and 3,794 human-verified atomic assertions (≈15.4 per video). Each atomic assertion is a standalone physical statement (e.g., "the cup fractures upon impact") judged for precision (claims are grounded) and recall (annotated physics is covered), combined into an F1 score.
Self-Evolving Caption Prompt
A fixed captioning model generates descriptions on 20 development videos. A physics evaluator flags omissions and errors; an evolution agent revises the shared caption prompt (not per-video patches). The updated prompt is validated on PhysCapBench. Over 9 rounds, PhysCapBench F1 rose from 78.64 to 87.82, with a temporary dip at round 2 (76.28) attributed to over-cautious descriptions that sacrificed recall. Crucially, the captioning model weights remain frozen throughout.
With a frozen Cosmos3-Nano generator, updating only the inference-time captions lifted PhyGenBench scores from 64.17 to 67.29, confirming that caption improvements transfer to generation quality.
Language-Guided Data Retrieval
Physis-Lang converts model weaknesses (e.g., deformation, collision, fracture errors) into physics-domain tags and retrieves real videos exhibiting those mechanisms, regardless of visual appearance. The final training set contains 183K real videos: 71K from WISA-80K plus 112K retrieved via this process, all re-captioned with the evolved prompt.
On VideoPhy-2, category-level gains include: chemical processes +8.00 pp, thermal processes +8.00 pp, fracture mechanics +7.45 pp, cloth deformation +7.19 pp.
Benchmark Results
Physics-IQ Verified: Cosmos3-Super 48.2, Cosmos3-Nano 43.3 (previous best 42.7); Super leads by 5.5 pp. The 16B Nano with Physis-Lang outperforms the original 64B Super.
Four physics benchmarks: Physis-Lang (Cosmos3-Nano) achieves open-source SOTA on all four; on three of them it surpasses Veo 3.1.
Evaluated on both image-to-video and text-to-video tasks against Veo 3.1, HunyuanVideo-1.5, Wan2.2, and physics-specialized methods.
Qualitative Examples
Generated samples demonstrate contact/collision (wood blocks toppling), tearing (paper pulled apart), and deformation under load (kettlebell pressing into a cushion).
The work argues that improving the language representation fed to world models is a viable path to better physical understanding and generation.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
