WorldRoamBench: Benchmarking Long-Horizon Stability in Interactive World Models

Amap, Nanjing University, Tsinghua, and Peking University release WorldRoamBench, a benchmark with 1000+ long-horizon roaming samples evaluating 12 interactive world models on action following, visual stability, physics adherence, and memory retention, revealing that even top models like Genie 3 show significant weaknesses in specific dimensions.

Machine Heart
Machine Heart
Machine Heart
WorldRoamBench: Benchmarking Long-Horizon Stability in Interactive World Models

Introduction: The Need for Long-Horizon Evaluation

The release of Genie 3 marks a shift from passive video generation to interactive world models where users input actions to change a continuously evolving world. However, existing evaluations remain limited to 5–10 second clips, failing to expose problems that emerge only during minute-long interactions, such as gradual visual degradation, action misalignment, physics violations, and memory loss.

WorldRoamBench: A Four-Dimensional Benchmark

WorldRoamBench, developed by Amap Visual Technology Center in collaboration with Nanjing University, Tsinghua University, and Peking University, provides 1000+ long-horizon roaming samples to systematically evaluate 12 mainstream interactive world models (including Genie 3, Happy Oyster, LingBot-World) across four capability dimensions: action, visual, physics, and memory. The benchmark covers first-person and third-person perspectives across natural, urban, and indoor scenes, with differentiated task designs to minimize cross-dimension interference.

Dimension 1: Action Following – Trajectory Correctness Masks Step-Level Errors

Final trajectory accuracy does not guarantee that every user command was executed correctly. Models exhibit varying action magnitudes, and trajectory-level metrics can hide local delays, misresponses, and direction errors. For interactive world models, timely, accurate, and continuous feedback to each input matters more than the final position. The article shows an example where the user continuously inputs "forward" commands, yet the generated video repeatedly moves backward mid-sequence, while the overall trajectory still appears correct.

Action following error: forward command causes backward movement, 3x speed
Action following error: forward command causes backward movement, 3x speed

Dimension 2: Visual Stability – Tracking Drift Over Time

Long interactions require sustained visual stability, but current evaluations often rely on average quality or first/last frame comparisons, missing mid-sequence drift. Errors accumulate over time, causing drift, deformation, or structural collapse. Even if quality later recovers, these events are diluted in average scores or invisible in endpoint comparisons. The article demonstrates a case where obvious visual drift occurs mid-sequence then recovers; comparing only first and last frames would completely miss the drift.

Visual drift mid-sequence with later recovery, 2x speed
Visual drift mid-sequence with later recovery, 2x speed

Dimension 3: Physics Adherence – Testing Collision Response

Physical realism must be verified during interaction, not just in static scenes. WorldRoamBench designs specific interaction tests that first confirm an interaction occurred, then judge whether the world's response follows physical laws. Coverage includes mechanics, optics, and 3D spatial consistency. An example shows a character walking toward a wall; instead of stopping, the model lets the character pass through the wall, revealing a physics violation only exposed by triggering a collision.

Physics violation: character walks through wall
Physics violation: character walks through wall

Dimension 4: Memory Retention – The "Death Rotation" Test

Memory is tested via a "death rotation": the model rotates, then returns along the opposite path to check scene consistency. Existing methods directly compare before/after frames, but differences may stem from position deviation rather than forgetting. WorldRoamBench first excludes position deviation, then measures how much of the original scene is retained and how much hallucinated content appears. In the demonstrated case, after returning with slight position offset, a white house that should be visible disappears entirely, indicating the model failed to retain the previously generated scene.

Memory test: white house disappears after return, 2x speed
Memory test: white house disappears after return, 2x speed

Evaluation Results: No Model Is an All-Rounder

Across 1000+ open-world long-horizon samples, no single model dominates all dimensions. On the first-person leaderboard, Genie 3, Lyra 2.0, and Happy Oyster rank top three; on the third-person leaderboard, Happy Oyster leads with Genie 3 second. However, Genie 3 exhibits clear "subject bias": it ranks first in memory and physics for first-person scenes, but only 8th in action following and 9th in visual quality.

First-person leaderboard rankings
First-person leaderboard rankings
Third-person leaderboard rankings
Third-person leaderboard rankings

Open Collaboration: Evolving Benchmark with Models

WorldRoamBench is not limited to the initial 12 models. The benchmark has opened an evaluation portal (https://worldroam.amap.com/) supporting test-set download and model inference result submission for scoring. Evaluation code and execution scripts will be released progressively. This marks a new phase for interactive world models: competition moves from short-term action following toward comprehensive long-horizon stability.

Paper: WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models Leaderboard & submission: https://worldroam.amap.com/
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

physics adherencememory retentionlong-horizon evaluationaction followingGenie 3interactive world modelsvisual stabilityWorldRoamBench
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.