WorldRoamBench: Benchmarking Long-Horizon Stability in Interactive World Models
Amap, Nanjing University, Tsinghua, and Peking University release WorldRoamBench, a benchmark with 1000+ long-horizon roaming samples evaluating 12 interactive world models on action following, visual stability, physics adherence, and memory retention, revealing that even top models like Genie 3 show significant weaknesses in specific dimensions.
Introduction: The Need for Long-Horizon Evaluation
The release of Genie 3 marks a shift from passive video generation to interactive world models where users input actions to change a continuously evolving world. However, existing evaluations remain limited to 5–10 second clips, failing to expose problems that emerge only during minute-long interactions, such as gradual visual degradation, action misalignment, physics violations, and memory loss.
WorldRoamBench: A Four-Dimensional Benchmark
WorldRoamBench, developed by Amap Visual Technology Center in collaboration with Nanjing University, Tsinghua University, and Peking University, provides 1000+ long-horizon roaming samples to systematically evaluate 12 mainstream interactive world models (including Genie 3, Happy Oyster, LingBot-World) across four capability dimensions: action, visual, physics, and memory. The benchmark covers first-person and third-person perspectives across natural, urban, and indoor scenes, with differentiated task designs to minimize cross-dimension interference.
Dimension 1: Action Following – Trajectory Correctness Masks Step-Level Errors
Final trajectory accuracy does not guarantee that every user command was executed correctly. Models exhibit varying action magnitudes, and trajectory-level metrics can hide local delays, misresponses, and direction errors. For interactive world models, timely, accurate, and continuous feedback to each input matters more than the final position. The article shows an example where the user continuously inputs "forward" commands, yet the generated video repeatedly moves backward mid-sequence, while the overall trajectory still appears correct.
Dimension 2: Visual Stability – Tracking Drift Over Time
Long interactions require sustained visual stability, but current evaluations often rely on average quality or first/last frame comparisons, missing mid-sequence drift. Errors accumulate over time, causing drift, deformation, or structural collapse. Even if quality later recovers, these events are diluted in average scores or invisible in endpoint comparisons. The article demonstrates a case where obvious visual drift occurs mid-sequence then recovers; comparing only first and last frames would completely miss the drift.
Dimension 3: Physics Adherence – Testing Collision Response
Physical realism must be verified during interaction, not just in static scenes. WorldRoamBench designs specific interaction tests that first confirm an interaction occurred, then judge whether the world's response follows physical laws. Coverage includes mechanics, optics, and 3D spatial consistency. An example shows a character walking toward a wall; instead of stopping, the model lets the character pass through the wall, revealing a physics violation only exposed by triggering a collision.
Dimension 4: Memory Retention – The "Death Rotation" Test
Memory is tested via a "death rotation": the model rotates, then returns along the opposite path to check scene consistency. Existing methods directly compare before/after frames, but differences may stem from position deviation rather than forgetting. WorldRoamBench first excludes position deviation, then measures how much of the original scene is retained and how much hallucinated content appears. In the demonstrated case, after returning with slight position offset, a white house that should be visible disappears entirely, indicating the model failed to retain the previously generated scene.
Evaluation Results: No Model Is an All-Rounder
Across 1000+ open-world long-horizon samples, no single model dominates all dimensions. On the first-person leaderboard, Genie 3, Lyra 2.0, and Happy Oyster rank top three; on the third-person leaderboard, Happy Oyster leads with Genie 3 second. However, Genie 3 exhibits clear "subject bias": it ranks first in memory and physics for first-person scenes, but only 8th in action following and 9th in visual quality.
Open Collaboration: Evolving Benchmark with Models
WorldRoamBench is not limited to the initial 12 models. The benchmark has opened an evaluation portal (https://worldroam.amap.com/) supporting test-set download and model inference result submission for scoring. Evaluation code and execution scripts will be released progressively. This marks a new phase for interactive world models: competition moves from short-term action following toward comprehensive long-horizon stability.
Paper: WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models Leaderboard & submission: https://worldroam.amap.com/
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
