NVIDIA Physis-Lang: Language-Driven Physics for Realistic Video Generation

NVIDIA, MIT, and Oxford researchers propose Physis-Lang, a framework that embeds detailed physical reasoning into language captions for video generation, using a self-evolving benchmark and language-guided data retrieval to achieve state-of-the-art results on physics benchmarks, surpassing Veo 3.1 on three of four tests.

Machine Heart
Machine Heart
Machine Heart
NVIDIA Physis-Lang: Language-Driven Physics for Realistic Video Generation

Physis-Lang: Language as Physics Representation for Video Generation

Researchers from NVIDIA, MIT, and Oxford University introduce Physis-Lang, a framework that embeds detailed physical reasoning — causes, laws, and outcomes — into natural language captions to improve the physical realism of generated videos. The approach operates across three stages: data selection, model training, and inference.

Physics-Aware Captioning and PhysCapBench

Standard video captions describe scenes and actions but omit the physical processes linking cause and effect. Physis-Lang enriches captions with continuous state changes. For example, instead of "butter melts," a Physis-Lang caption reads: "Continuous heating gradually melts the butter; the solid portion collapses under gravity, shrinking in volume, while the melted liquid forms a shallow pool that expands outward." This forces the model to represent temperature-driven phase transitions, gravity-driven deformation, and concurrent solid–liquid dynamics.

To evaluate caption quality, the team built PhysCapBench , comprising 246 physics videos and 3,794 human-verified atomic assertions (≈15.4 per video). Each atomic assertion is a standalone physical statement (e.g., "the cup fractures upon impact") judged for precision (claims are grounded) and recall (annotated physics is covered), combined into an F1 score.

Self-Evolving Caption Prompt

A fixed captioning model generates descriptions on 20 development videos. A physics evaluator flags omissions and errors; an evolution agent revises the shared caption prompt (not per-video patches). The updated prompt is validated on PhysCapBench. Over 9 rounds, PhysCapBench F1 rose from 78.64 to 87.82, with a temporary dip at round 2 (76.28) attributed to over-cautious descriptions that sacrificed recall. Crucially, the captioning model weights remain frozen throughout.

With a frozen Cosmos3-Nano generator, updating only the inference-time captions lifted PhyGenBench scores from 64.17 to 67.29, confirming that caption improvements transfer to generation quality.

Language-Guided Data Retrieval

Physis-Lang converts model weaknesses (e.g., deformation, collision, fracture errors) into physics-domain tags and retrieves real videos exhibiting those mechanisms, regardless of visual appearance. The final training set contains 183K real videos: 71K from WISA-80K plus 112K retrieved via this process, all re-captioned with the evolved prompt.

On VideoPhy-2, category-level gains include: chemical processes +8.00 pp, thermal processes +8.00 pp, fracture mechanics +7.45 pp, cloth deformation +7.19 pp.

Benchmark Results

Physics-IQ Verified: Cosmos3-Super 48.2, Cosmos3-Nano 43.3 (previous best 42.7); Super leads by 5.5 pp. The 16B Nano with Physis-Lang outperforms the original 64B Super.

Four physics benchmarks: Physis-Lang (Cosmos3-Nano) achieves open-source SOTA on all four; on three of them it surpasses Veo 3.1.

Evaluated on both image-to-video and text-to-video tasks against Veo 3.1, HunyuanVideo-1.5, Wan2.2, and physics-specialized methods.

Qualitative Examples

Generated samples demonstrate contact/collision (wood blocks toppling), tearing (paper pulled apart), and deformation under load (kettlebell pressing into a cushion).

The work argues that improving the language representation fed to world models is a viable path to better physical understanding and generation.

Physics-IQ Verified leaderboard comparison
Physics-IQ Verified leaderboard comparison
Self-evolution gains on PhysCapBench and PhyGenBench
Self-evolution gains on PhysCapBench and PhyGenBench
Language-guided data retrieval flow
Language-guided data retrieval flow
Per-category VideoPhy-2 improvements
Per-category VideoPhy-2 improvements
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

video generationNVIDIAphysics simulationlanguage representationPhysCapBenchPhysis-Lang
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.