Behavior Consistency Beats State Consistency in Text World Models for Agents
The paper introduces BehR, a behavior consistency reward for training text-based world models, showing that optimizing for agent decision alignment rather than text fidelity improves trajectory-level consistency across 16 configurations, reduces false positives in offline evaluation from 42.5% to 9.5%, and enhances lookahead planning for weaker agents.
