Agentic Visual Generation: A Five-Level Hierarchy of Controller Decision-Making

This survey introduces a five-level hierarchy (L0–L4) for agentic visual generation systems, classifying them by how far the controller can influence generation decisions—from fixed support to cross-task experience reuse—and proposes a level-conditioned evaluation protocol to isolate causal contributions of control scope.

Data Party THU
Data Party THU
Data Party THU
Agentic Visual Generation: A Five-Level Hierarchy of Controller Decision-Making

Introduction

Visual generation has expanded from images to video, 3D, slides, user interfaces, and world modeling. Models such as DALL-E 2, Imagen, Parti, Latent Diffusion, SDXL, DALL-E 3 improve image quality and language alignment; Video Diffusion Models, Make-A-Video, Imagen Video, Video LDM, Lumiere extend diffusion to the temporal dimension. These advances strengthen the executor —the ability to produce better content given a condition.

Agentic visual generation depends not on a stronger generator but on a controller that can influence the generation process. The controller, typically an LLM, VLM, or multimodal model, rewrites prompts, constructs layouts, selects models, invokes editors, verifies intermediate results, decides retries, and retains experience to alter future routing. Current literature treats planning depth, tool use, multi-agent collaboration, and reinforcement learning as evidence of agency, but none is a decisive criterion. The decisive question: which generation decisions can the controller directly control? Can it change the next step after observing a result? Can it transfer experience from one task to subsequent independent tasks?

From Visual Generation to Generation-Level Control

Core Concepts

Visual generation describes what the executor can create or edit. Generation-level control describes how the controller chooses conditions, invokes operations, and changes subsequent actions around the executor. The controller can be a language model, multimodal model, learned policy, search process, or a team of specialized roles; its defining feature is intermediate decisions that affect task execution. Such decisions include prompt revision, layout construction, model routing, reference selection, editing, verification, memory update, and stopping. The generator constructs visual state—images, video, visual stories, slides, UIs, 3D scenes, or interactive environments. For slides and UIs, action representations may be PowerPoint objects, XML, HTML, CSS, or executable code, but the evaluation target remains the rendered visual structure and behavior.

System Definition

An agentic visual generation system has both a visual generator/editor and a decision process that controls generation over one or more steps. Given a user goal, current multimodal state, actions, and subsequent observations, the controller selects actions based on goal and state; the environment state is updated by generated images, video clips, scenes, critiques, or retrieved references. This formalization emphasizes three points: (1) the system creates part of its own future observation space because generation results become subsequent state; (2) the action space can mix symbolic actions, tool calls, and visual generation; (3) trajectory quality depends not only on final output but also on decisions made to obtain that output.

Task Axis and Mechanism Axis

Two auxiliary axes organize the landscape. The task axis covers image generation/editing, video generation/editing, slide and UI generation, 3D asset/scene construction, world-grounded synthesis, and interactive simulation. Task type does not determine agentic level—an image system may be more adaptive than a video system, and a multi-shot video pipeline may still be open-loop. The mechanism axis covers intent grounding, explicit/implicit planning, tool routing, retrieval, multi-agent collaboration, verification, memory, and reinforcement learning. These explain how the controller is implemented but do not define levels; tool use is merely an action space, multi-agent a topology, RL an optimization method—all can appear at different levels.

A Hierarchy of Controller Decision-Making Scope (L0–L4)

Level Assignment Principle

The core organizing principle is causal: agency in visual generation depends on how deep in the generation trajectory the controller can change future generation decisions. Conditions occur before execution; execution decisions determine which visual operation is called; observing results can redirect subsequent actions within the current task; persistent experience can influence a different task after the current one ends. This yields five labels:

L0 Fixed Support : fixed components (generators, editors, retrievers, evaluators, reward models, benchmarks, preset pipelines) that have no deployment-time generation-level decisions.

L1 Conditioning Control : controller constructs input conditions for a fixed executor but does not choose specific visual operations.

L2 Execution Control : controller selects and invokes generation, editing, rendering, or other content-modifying operations.

L3 Outcome-Adaptive Control : controller changes subsequent actions within the current task based on intermediate results.

L4 Experience-Adaptive Control : controller retains experience from completed tasks and alters decisions for future independent tasks.

Conservative Assignment Rule

Judge from highest to lowest: if cross-task persistent experience influences future control → L4; if observing results changes subsequent actions in current task → L3; if controller selects and invokes visual operations → L2; if only constructs conditions for fixed executor → L1; otherwise → L0. This rule handles boundary cases: prompt rewriting + fixed generator = L1; generator/editor selection = L2; critique + repair/regeneration = L3; retaining only current-task state ≠ L4; abstracting reusable skills from completed tasks = L4; RL-trained fixed generator may still be L0/L1 because level depends on inference-time action space, not training optimizer.

Fine-Grained Classification

Within each level, a secondary axis organizes systems: L0 by support function; L1 by specification for the generator; L2 by executable object the controller chooses; L3 by feedback source triggering next action; L4 by persistent experience carrier. This avoids inventing new tags per paper and stabilizes cross-modal comparison.

Hierarchy diagram showing L0-L4 levels
Hierarchy diagram showing L0-L4 levels
Fine-grained classification within each level
Fine-grained classification within each level

L0: Fixed Support

Generation, Retrieval, and Training Infrastructure

L0 includes fixed generators, editors, retrievers, evaluators, reward models, benchmarks, and preset pipelines (e.g., Latent Diffusion, Video Diffusion, DreamFusion, ControlNet, T2I-Adapter, DreamBooth, IP-Adapter, InstructPix2Pix). They can be important components of later agentic systems but have no deployment-time generation-level decisions. Generator controllability ≠ controller decision scope: ControlNet executes given spatial conditions, IP-Adapter accepts reference images, InstructPix2Pix follows edit instructions, but if the system does not decide whether to call another operation, repair a failed result, or retain experience, it remains fixed support.

Data construction and training infrastructure also belong to L0. Works like ScaleEdit-12M, JAVEDIT, DataEvolver, AgentComp, Gen-n-Val may use multi-role or tool pipelines in offline data construction; DPOK, AeSlides, FrontCoder, T2I-R1, ReasonGen-R1 improve generators via reward, reasoning supervision, or preference optimization. But if the deployed inference path is fixed, they do not constitute generation-level agentic control.

Evaluators and Boundary Tests

Evaluators and benchmarks (ImageReward, Pick-a-Pic, GenEval, T2I-CompBench, VBench, EvalCrafter, AgenticVBench, DECKBench, 3DCodeBench) diagnose generation quality, compositional failures, video temporal defects, slide/3D modeling issues. They become part of an L3 closed loop only when mapped by another mechanism to subsequent repair, re-routing, or stop actions. The counterfactual test for L0: given same goal and state, can the deployed system choose a different generation action? If not → L0; if only conditions for fixed executor → L1; if selects operation → L2; if changes subsequent actions based on result → L3; if retains cross-task experience → L4.

L0 boundary test illustration
L0 boundary test illustration

L1: Conditioning Control

Specifications for Fixed Executors

L1 takes the first step from L0's fixed path: the controller converts the goal into executor-consumable conditions. Formally, it generates conditions, then the fixed generator produces the result. Control stops there; the system commits before observing output. L1 divides into five categories:

Text prompt specifications (Promptist, TIPO, APE, DiffChat, POSI) rewrite user requests into prompts better suited for the generator while preserving semantics and safety.

Spatial and geometric specifications (LMD, LayoutGPT, LLMControl, NaLA, DAC-Pose) make object boxes, regions, poses, or scene structures explicit.

Retrieval evidence specifications (World-To-Image, Gen-Searcher, Cross-modal RAG, RealRAG, MosAIG) decide whether to retrieve, what to retrieve, and how to bind evidence to concepts or attributes.

Temporal and camera specifications (Aurora, TempAct, AgentHOI, VideoGen-of-Thought, ShotVerse, CinemaTraj) commit to actions, shots, identities, and viewpoints before video/scene generation.

Structured content specifications (CANVAS, S2ED, MangaFlow, SlideTailor, ScreenCoder, PosterGen) output storyboards, panels, documents, or UI structures.

Primary Failure Modes

L1's core benefit is reducing ambiguity, but it binds the system more tightly to pre-generation decisions. Prompt optimization may improve aesthetics yet distort user intent; spatial planning may miss entities or encode unrealizable layouts; retrieval systems may introduce irrelevant evidence that the generator faithfully follows. Therefore L1 cannot be judged solely by final image preference scores. Prompt rewriting must test semantic preservation; layout planning must test geometric validity; retrieval augmentation must test evidence relevance. Natural language prompts are portable but weakly constraining; boxes, trajectories, and scene descriptions constrain more strongly but depend on generator support for corresponding control formats.

L2: Execution Control

From Specifications to Executable Actions

L2 shifts the action interface from "given specification" to "select and invoke operation." The controller can choose generators, editors, search tools, renderers, programs, workflow graphs, or domain-specific languages. Visual ChatGPT exposes visual foundation models as callable tools; ComfyUI-Copilot constructs executable node graphs; ViMax organizes screenwriting, shot planning, character style, and clip generation into a video operation chain. L2 categorizes by primary executable object: model/tool operations, image/structured graphics operations, video/audiovisual operations, document/UI operations, 3D/CAD/world operations. Workflow construction and multi-role collaboration are merely implementation strategies within these action spaces.

Open-Loop Boundary

L2 typically remains open-loop: the system can choose a route, but the generation result does not necessarily change the next route. A long plan, many roles, a complex executable workflow do not automatically imply feedback control. Only when execution returns observations that affect subsequent generation-level actions does the system enter L3. This demands evaluation reporting plan validity, tool parameters, execution success rate, state maintenance, and output quality separately. A system may get a good image by chance route, or have a correct route but generator failure. Separating trajectory correctness from terminal quality reveals whether the controller truly contributes.

L3: Outcome-Adaptive Control

Feedback Sources

L3 adds result-dependent state transitions: after an action produces an observation, the controller updates state, then selects the next action. L3 classifies by the key observation source that triggers subsequent actions:

Perceptual result feedback : visible discrepancies in rendered images, video clips, multi-view projections, slides, or UI screenshots change subsequent actions (e.g., RPG, SLD, GenPilot, GenArtist, PPTAgent, UI2CodeN).

Structured and execution feedback : program state, workflow state, timeline state, source material state, scene or engine state.

Physical and constraint feedback : simulation dynamics, collisions, stability, reachability, world rule checks.

Human review feedback : only when explicit human judgment changes subsequent actions within the current task.

L3 feedback loop diagram
L3 feedback loop diagram

Diagnosis, Maintenance, and Stopping

The difficulty in L3 is not generating a few more times but correctly coupling diagnosis and action. Prompt revision translates visual defects back to language; tool editing can choose local operations; code/XML/scene state repair can pinpoint specific objects. The latter two aid credit assignment but require more precise critics. Maintenance and stopping are equally critical: a local repair that breaks previously satisfied constraints is not progress. Reporting should include non-target changes, regression of accepted content, gain per revision, and whether stopping is calibrated. Video, UI, slides, and 3D scenes especially need tracking of temporal state, executability, interaction behavior, and constraints in invisible viewpoints.

L4: Experience-Adaptive Control

Five Persistent Experience Carriers

L4's defining capability is crossing the current task boundary: experience from completed tasks is written into a carrier and influences future independent tasks. Five carrier types:

Capability and tool profiles (DiffusionAgent, OctoT2I, PerfGuard, GenRouter) update future model/tool routing based on historical quality, efficiency, or ranking.

Episodic and user memory (MemoGen, BrandFusion, UniVA, MemSlides) retrieve past requests, results, or user preferences to alter new task control.

Reusable processes and skills (GenEvolve, EvoDiagram, ManimAgent) abstract readable strategies from success/failure trajectories.

Executable workflows and control frameworks (COMFYCLAW, FigAgent, AutoDesign, VideoWeaver) upgrade graphs, programs, or middleware into verifiable, versionable persistent assets.

Policy and model updates (SIDiffAgent, JarvisEvo, SPIRAL, CLARE) write experience into parameters or policies so future control needs no explicit memory retrieval.

Transfer, Failure, and Rollback

L4 is most easily overestimated because experience can preserve useful strategies but also harmful biases. Temporal hold-out task evaluation is required for L4, reporting positive transfer, negative transfer, forgetting, stale experience sensitivity, provenance tracking, and rollback capability. An "self-evolving" label without cross-task behavioral evidence does not prove L4. Reversibility varies greatly: profile entries and episodic memories can be deleted; processes can be revised; executable workflows can be versioned; but harmful parameter updates may require rollback checkpoints or retraining. Thus L4's true value must measure both future task gains and the cost of detecting and undoing bad updates.

Training and Reinforcement Learning for Generation Controllers

Training Does Not Determine Level

Level describes what the controller can change at inference time; training describes how those capabilities are learned. Supervised fine-tuning or RL does not determine system level. A fixed generator optimized with image reward may remain L0; a supervised-trained router can be L2; an untrained repair loop can be L3. Controller learning breaks into four stages: trajectory data recording decisions and consequences, supervised fine-tuning initializing behavior from demonstrations, feedback providing credit signals, policy optimization improving the policy.

Data, Rewards, and Policy

Generation controller data should preserve request, state, action, observation, resource usage, and termination—not just final images. Failure and repair trajectories are especially important because they reveal whether errors stem from prompt, retrieval, routing, masking, generator, verifier, or stopping rule. Long-horizon and cross-task systems also need tool versions, timestamps, provenance, and rollback info. Rewards must distinguish terminal preference from process verification. ImageReward, Pick-a-Pic suit overall preference but cannot judge which step caused success. Process rewards should check tool execution, constraint satisfaction, diagnosis-repair consistency, accepted content preservation, and whether continuing is worth the cost. For real systems, rewards should explicitly deduct compute, tool calls, interaction, and monetary costs. Policy optimization targets may be fixed generator, conditioning policy, workflow controller, or cross-task memory/skill updates. Reporting optimization object, interaction mode, and persistence scope avoids mistaking strong generators, large search budgets, or reward overfitting for controller progress.

Evaluation and Benchmarking

Level-Conditioned Counterfactual Evaluation

The second major contribution is level-conditioned evaluation : constructing matched counterfactuals per level to isolate the causal value of wider control scope, rather than comparing systems with different tools, budgets, and generators.

L0 asks what the fixed executor can do: fix input conditions, random seeds, sampling parameters, candidate budget.

L1 asks whether better specifications improve the same executor: fix generator and sampling budget, vary only conditioning policy.

L2 asks whether the controller chose a better route: fix available models, tools, action budget, evaluator; compare fixed, random, frequency, or oracle routing.

L3 asks whether result feedback leads to successful repair: compare open-loop vs. closed-loop with/without intermediate result access.

L4 asks whether historical experience improves future tasks: remove, shuffle, replace with stale or irrelevant experience.

Level-conditioned evaluation framework
Level-conditioned evaluation framework

Common Control Variables

Four categories of common evidence: (1) Decision validity : constraint coverage, dependency consistency, invalid call rate, parameter correctness, evidence relevance, memory relevance. (2) Efficiency and trajectory utility : call count, candidate count, edit operations, tokens, latency, compute time, monetary cost. (3) Robustness and recovery : ability to recover from invalid layouts, unavailable tools, corrupted intermediate results, critic errors, stale memory. (4) Faithful attribution : do planned entities correspond to rendered regions? do claimed actions match tool calls? does diagnosis predict successful repair? Iterative systems are prone to selection bias: generating eight candidates and reporting the best cannot be directly compared to single-sample baselines. Internal critics, stopping critics, and hold-out evaluators should be separated; human evaluation or independent verifiers must audit cases where internal scores rise but hard constraints fail.

Challenges and Future Directions

From Declarative to Executable Control

The L1→L2 challenge is making specification policies bear execution responsibility. The controller must not only write prompts or layouts but decide whether to call retrieval, editors, 3D assets, code execution, or other tools. This expands safety, privacy, licensing, and provenance risks. Future systems should attach references, models, licenses, parameters, and edit logs in machine-readable form to trajectories and perform permission-aware routing before execution.

From Execution to Reliable Feedback Control

The L2→L3 challenge is obtaining sufficiently reliable observations and mapping them to correct repairs. Future feedback should combine local multimodal diagnosis, specialized detectors, executable checks, simulation state, uncertainty, and human escalation. Diagnosis must map to actionable regions, objects, frame intervals, or code components; otherwise repeated generation merely boosts selector scores without fixing real defects.

From Current Task State to Reusable Experience

The L3→L4 challenge is compressing current task state into future-usable, auditable experience. Tool prices, interfaces, user preferences, and safety policies change; old experience may become harmful. Persistent memory therefore needs timestamps, provenance, conflict detection, selective forgetting, and rollback. Long video, 3D scenes, and visual stories may retain long-term state but remain L3; only when it changes new task control after project end does it become L4.

Generator as Controller

A potential new paradigm: generator-as-controller . Current systems use LLM/VLM to interpret goals and select actions while visual generators only execute external conditions. This creates an interface bottleneck: visual state is rendered, encoded into language, then converted back to prompts or control signals. If the generator itself also controls, it can act on the same state that produces visual content, linking defects directly to visual tokens, latent regions, frame intervals, scene objects, or generation steps—shortening feedback loops, reducing interaction cost, improving local repair and credit assignment. Even so, a unified model does not automatically equal agency. It must prove its visual state leads to different subsequent actions and that persistent experience changes future tasks; otherwise it remains a fixed generation path.

Conclusion

The survey's greatest value is freeing agentic visual generation from buzzwords and grounding it in the decidable question: "Which generation decisions can the controller change?" L0–L4 are ordered not by model size, output quality, tool count, or multi-agent usage, but by the farthest causal reach of the controller: from fixed support, to condition construction, execution selection, result repair, to cross-task experience reuse. Corpus statistics show recent growth concentrated in L3 (using intermediate results for within-task repair); truly mature L4 cross-task experience remains scarce. The next phase's key is not just stronger generative models but executable action interfaces, reliable visual verification, causal credit assignment, transparent resource accounting, selective long-term memory, editable collaborative state, and auditable provenance records. Agentic visual generation represents not "image generation extended to video or 3D" but a change in computational unit: from single conditional sampling to an observable, repairable, memorable, evaluable generation trajectory unfolding around a goal. Future generator-as-controller may unify generation, observation, tool selection, trajectory revision, and experience reuse in a single visual policy, but it must still pass level-conditioned counterfactual evaluation proving that wider control scope yields attributable gains.

Paper: https://arxiv.org/abs/2609.06758
PDF: https://arxiv.org/pdf/2609.06758

Code example

来源:专知
本文
约7000字
,建议阅读
15
分钟
这篇综述最有价值的地方,是把智能体式视觉生成从一组流行关键词中解放出来,重新落到“控制器能改变哪些生成决策”这一可判定问题上。
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

reinforcement learningagentic visual generationcontroller hierarchycross-task experiencegeneration-level controlL0-L4 taxonomylevel-conditioned evaluationvisual generation agents
Data Party THU
Written by

Data Party THU

Official platform of Tsinghua Big Data Research Center, sharing the team's latest research, teaching updates, and big data news.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.