H-JEPA: LeCun's Hierarchical World Model for Visual Planning, Open-Sourced
LeCun's startup AMI releases H-JEPA, a hierarchical world model that stacks JEPA layers to plan at multiple time scales, achieving 73% success on Visual AntMaze versus 18% for single-layer models, with code and pre-trained weights fully open-sourced.
LeCun's Startup AMI Releases H-JEPA: Hierarchical World Model for Visual Planning
LeCun's Advanced Machine Intelligence (AMI) has published the H-JEPA paper, co-authored with researchers from NYU, INRIA Paris, and Brown University. H-JEPA introduces a hierarchical world model that stacks multiple JEPA (Joint Embedding Predictive Architecture) layers, each predicting future states at different time scales. The model is fully open-sourced with code and pre-trained weights.
Why Hierarchical Planning?
LeCun has advocated hierarchical JEPA since 2022: models should predict at different abstraction levels and time scales, then use hierarchical planning to decompose complex goals into actions. The article illustrates with a travel analogy: planning a trip from New York to Paris involves high-level decisions (flights, dates), mid-level navigation (leaving office, taxi to airport), and low-level motor control (walking down stairs, foot placement). Each level operates on different information and time horizons. Similarly, robots need high-level task understanding and low-level motor control; a quadruped robot may reach a target location but still have mismatched leg posture, causing a single-level model to misjudge proximity.
H-JEPA Architecture
Each H-JEPA layer contains three components:
State Encoder: Extracts environment representation. The lowest layer processes raw images; higher layers further abstract lower-layer representations, retaining information relevant at that time scale.
Action Encoder: Describes state changes caused by actions. Low layers encode concrete actions; higher layers compress action sequences into coarser changes.
Predictor: Combines current state and action to predict future state representations. Adjacent layers cover different time spans; in experiments, one high-layer step corresponds to two low-layer prediction steps.
Layers are trained end-to-end. High-layer predicted intermediate states become subgoals for the lower layer. The lower layer maps its predicted states into the high-layer representation space, compares them to the subgoal, and adjusts actions accordingly.
Preventing Representation Collapse
To avoid representation collapse (where the encoder maps all inputs to the same vector), H-JEPA adopts SIGReg from the earlier LeJEPA work, constraining the representation distribution so distinct states are not compressed into identical vectors.
Experimental Results
On the Visual AntMaze simulated maze task, a three-layer H-JEPA achieves a 73% success rate, compared to 18% for a single-layer model. The hierarchical model also achieves better performance with less planning computation across multiple simulated tasks. Analysis shows higher layers gradually discard leg posture details while preserving robot position information.
However, hierarchy has limits: on the Push-T task, a two-layer model performs well, but adding a third or fourth layer degrades performance. The authors hypothesize that shorter task trajectories limit the training data available for higher layers.
LeCun's Persistent Vision: JEPA Over LLMs
H-JEPA continues the core JEPA principle: prediction in representation space. Unlike LLMs that predict next tokens, visual JEPA encodes images/video into internal representations and predicts target representations without generating pixels. In world models, the predictor also conditions on actions to forecast future state representations; a planner then compares action consequences to find paths to goals.
If you're interested in human-level AI, don't work on LLMs.
LeCun reiterated this view at ECCV 2024 in Malmö. When questioned about the slide's authenticity, he confirmed "The package is real." He acknowledges LLM-based tools are useful but maintains LLMs alone cannot reach human-level intelligence.
AMI Status and Competitor Landscape
AMI raised $1.03 billion at a $3.5 billion pre-money valuation in March 2024. The team includes LeCun (Chairman), Alexandre LeBrun (CEO), 谢赛宁 (Chief Scientist), Pascale Fung (Research & Innovation), and Michael Rabbat (World Models, co-author of H-JEPA). AMI's website currently showcases technical direction, team, and hiring, with no public product yet.
Meanwhile, World Labs (founded by Fei-Fei Li) agreed to an $82 billion all-stock acquisition by AMD on September 28, 2024. One world-model startup moves toward chip integration; the other continues its research trajectory.
References: VentureBeat article (https://venturebeat.com/technology/h-jepa-teaches-world-models-to-plan-at-multiple-levels-of-abstraction), arXiv preprint (https://arxiv.org/pdf/2610.06805).
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
