Structured Representations as Scaffolding for Robot Reasoning: Jiajun Wu's IROS 2026 Talk

Jiajun Wu's IROS 2026 talk presents structured representations as scaffolding for robot reasoning, covering learning predicates and actions from demonstrations, differentiable planning with foundation models, post-training for unseen objects, and generating interaction sequences via video diffusion models without human demonstrations.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Structured Representations as Scaffolding for Robot Reasoning: Jiajun Wu's IROS 2026 Talk

Stanford assistant professor Jiajun Wu delivered an invited talk at the RoBoWoMo workshop during IROS 2026 titled "Building Physical Agents via Structured Representations." The talk traces a research arc from visual scene understanding to building robots that can reason and act in the physical world.

The Kitchen Scenario: Why Structured Reasoning Is Hard

Wu opens with a simple kitchen scene: a plate, a faucet, a dab of blue dish soap. A student starts washing the plate, then leaves. The robot must take over. The task seems straightforward — wash the plate and put it away — but the robot must first identify which object is the plate, whether it is dirty, whether soap is present, whether the faucet is already running, and whether the plate has already been washed. The same high-level instruction leads to completely different action sequences depending on the initial state.

Wu illustrates three distinct challenges:

Adapting to others' actions: A human washes the plate and puts it back; the robot must recognize the plate is clean and only needs to be stored, not washed again.

Subtle initial-state differences: The only visual difference is a tiny blue blob of soap. That small perceptual change causes a qualitative shift in the state representation: without soap, washing would make a mess; the robot must first apply soap, then wash, then store.

Updating with new observations: After placing the plate on the rack, the robot notices grime underneath. The goal "make everything clean" now implies wiping the counter, even though no one explicitly commanded it.

These variations demand a system that can maintain consistency across environments, objects, and goals.

A Paradigm: Structured Representations as a Scaffold

For several years Wu's group has pursued a paradigm: raw inputs (video or language) → foundation models discover compositional abstractions → abstractions are reused and recombined for general tasks → learned control policies execute actions. The goal is to learn from "natural supervision" — videos, demonstrations, or language — with minimal human engineering.

The architecture consists of two main learned modules:

Predicates: Neural networks that evaluate properties (dirty/clean) and spatial relations (2D/3D) for arbitrary object sets.

Atomic actions with preconditions and effects: Each action (e.g., "wash") specifies preconditions (object grasped, robot in kitchen) and effects (if object has soap, it becomes clean). A diffusion policy per action maps states to joint torques.

The symbolic layer defines compositionality: actions chain when preconditions are satisfied. Perception and control are fully neural.

Learning the Domain-Specific Language (DSL)

Early versions used hand-crafted or PDDL-derived DSLs. The key question: is a fixed DSL general enough? The next step lets a foundation model propose the DSL from data (demonstration videos, operation logs, or language). The model suggests relevant predicates, atomic actions, preconditions, and effects. Verification is twofold: symbolic consistency checks and empirical validation against the demonstration data (e.g., does every "wash" in the data indeed follow a "grasp"?).

Wu compares this to coding agents like Cursor: given a high-level request, the agent produces a multi-step plan, executes it, and verifies correctness. The robotics analogue moves this loop into physical space.

When Post-Training Is Necessary

Foundation models trained on internet-scale data lack knowledge of proprietary objects (e.g., a factory-specific part). When a predicate network cannot judge "dirty" on a never-before-seen object, post-training becomes useful. Because the entire chain — predicate → action → effect — is differentiable, the final execution outcome (the plate became clean) provides a supervision signal that can backpropagate to update the predicate network.

Learning to Plan with Foundation Models

Recent work replaces the traditional symbolic planner with a neural planner. The planner receives natural-language predicates and images as prompts to a foundation model, which outputs action sequences. Because open-source vision-language models exhibit weaker visual reasoning, the team fine-tunes them on robot data. The fine-tuned planner produces executable sequences with far fewer human interventions, and it conditions on the robot's specific skill library (e.g., knowing the faucet must be turned on before filling a mug).

Zero-Shot Interaction Generation Without Demonstrations

Moving beyond demonstration-dependent learning, Wu describes a system that, given a single static view of a novel object (a faucet, a box), samples a distribution of possible interaction modes (turn, open, close). Trained on diverse objects, the model generalizes to unseen categories (monitor, USB drive, flower) and even to noisy iPhone scans of entire rooms (dishwasher, fridge, microwave). The generated interactions are not memorized; they reflect the learned distribution of "how can I act on this?"

To make this goal-directed, language conditioning is added: "a child jumps on a seesaw" generates a dynamic 4D scene showing the interaction. The method leverages video diffusion models and Score Distillation Sampling (SDS) extended to 4D: a static 3D representation is lifted to 4D (adding time), rendered from multiple views, and the video diffusion model guides the 4D representation toward realistic motion.

Concrete results include:

Generating a robot picking up a brick conditioned on "robot picking up brick."

Generating a human unfolding a cloth, providing dense 3D point trajectories as free supervision for a robot policy that then executes the unfolding in the real world.

Generating a laptop closing sequence.

Speed improvements via caching and semantic rendering enable longer-horizon sequences, yielding not just atomic skill demonstrations but complex, multi-task demonstration data for skill training.

Conclusion: Structure Is Scaffolding, Not the Answer

Wu emphasizes that structured representations are not a rigid framework to be hard-coded; they are an external scaffold that helps the system learn more efficiently and enables human communication. The ultimate aim is to use data-rich modalities (video, language, 3D/4D) to build reasoning capabilities: atomic representations (wash, place) plus their dependencies and effects enable generalized long-horizon problem solving.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

task planningfoundation modelszero-shot generalizationdiffusion policyvideo diffusion modelsstructured representationsIROS 2026Jiajun Wupredicate learningrobot reasoning
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.