VideoGen-Agent: 8B Agent Autonomously Calls Tools to Boost Video Generation
VideoGen-Agent, a Qwen3-VL-8B-based multimodal agent, learns via supervised fine-tuning and reinforcement learning to autonomously invoke retrieval, simulation, and verification tools for video generation, scoring 75.6 on VABench (19.1 points above base) and 86.1 with upgraded tools, while humans prefer its outputs 84.3% over Seedance 2.0.
Overview
VideoGen-Agent reframes video generation as a learnable decision process. A multimodal agent based on Qwen3-VL-8B-Instruct actively retrieves knowledge, fetches visual references, runs physics simulations, and verifies generated clips before final output. The agent is trained first with supervised fine-tuning (SFT) on 16K teacher trajectories, then with multi-task reinforcement learning (RL) on 8K held-out prompts, using a mixed reward that weights call validity (0.1), task-appropriate tool use (0.4), and video quality (0.5).
Three Illustrative Examples
1. Procedural knowledge – Tai Chi move "White Crane Spreads Wings" : The base model produced a generic arm raise. The agent first retrieved textual descriptions of the move (weight shift, right knee bend, left foot forward) and fed those details to the generator, yielding keyframes that show the distinct high-low hand posture.
2. Subject identity – Game character holding a crossbow : The base model generated a realistic hooded figure but missed the character's specific visual style. The agent retrieved reference images of the character and passed them to a conditional video generator, producing frames that match the reference's hairstyle, outfit, and overall aesthetic.
3. Temporal structure – 12-second video in three sequential shots : The prompt specified three phases (stressed gutter, cracking bracket, gutter swinging down). The agent split the timeline, generated each shot conditioned on the previous shot's last frame, and thereby satisfied each phase in order.
Agent Workflow
The agent operates in a ReAct loop: given the user prompt and interaction history, it selects among three tool groups – information augmentation (text retrieval, image retrieval, physics simulation), video generation (various conditional generators), and result verification (object detection, depth estimation). A system prompt defines tool schemas and typical patterns; the policy decides which tool to call, with what parameters, and whether to continue verifying or generating. Each tool's output becomes the next step's observation.
Training: From Imitation to Reinforcement
Supervised Fine-Tuning (16K trajectories) : Teacher demonstrations show how to choose tools, fill parameters, and use returned information.
Reinforcement Learning (8K prompts) : The agent resamples its own trajectories. Rewards combine:
Call validity (format correctness) – weight 0.1
Tool-use relevance (whether the right tool was chosen and its result utilized) – weight 0.4
Video quality (task-specific criteria, scored by Gemini 3.1 Pro) – weight 0.5
Task-balanced advantage normalization : Per-category rewards are normalized so that learning signals across the six task types are comparable.
VABench: Benchmark for Video Generation Agents
VABench contains 600 unseen prompts (100 per category) covering six capabilities:
Procedural Knowledge (PK) – professional actions/processes
Single Subject Identity (SI) – visual fidelity of a specified subject
Multiple Subject Identity (MI) – multiple specified subjects
Physical Consistency (PS) – collision and falling simulations
Scene Composition (CS) – spatial relations among objects
Multi-Shot Temporal Structure (MS) – 2-3 sequential shots, each ≤5 seconds
Prompts were filtered via deduplication, VLM screening, and human review. Evaluation uses Gemini 3.1 Pro with category-specific rubrics, normalized to 0-100; overall score is the unweighted mean of the six categories.
Main Results
Base Generator: 56.5 overall
VideoGen-Agent Toolset 1: 75.6 overall (+19.1 over base)
VideoGen-Agent Toolset 2 (upgraded generator): 86.1 overall
Toolset 2 improves over Toolset 1 across all categories (e.g., MI from 65.3 to 86.7, MS +25.7 over Seedance 2.0). Human side-by-side comparisons (400 judgments each) show:
Toolset 2 vs. Seedance 2.0: 84.3% preference for Toolset 2
Toolset 1 vs. Seedance 1.0 Pro Fast: 88.0% preference for Toolset 1
Auto-evaluator (Gemini 3.1 Pro) agrees with human majority vote on 92.0% of 600 video pairs (per-category: PK 72%, SI 98%, MI 100%, PS 94%, CS 88%, MS 100%).
Tool Upgrade Reusability
Swapping the generation backend from Toolset 1 to Toolset 2 (stronger video models) without retraining the agent yields consistent improvements, demonstrating that the learned organizational policy transfers across compatible tool interfaces. The largest gains appear in tasks that heavily rely on external knowledge, references, or explicit temporal structure (MI +21.4, MS +25.7, PK +14.1), while scene composition (CS) improves only marginally (+2.9), indicating some tasks are already well-handled by strong generators alone.
Ablation Studies
Using Toolset 1, the paper compares:
Prompt rewriting only : limited gain.
Fixed tool pipeline (hand-crafted ReAct flows): 64.1 overall.
SFT only : 69.2 overall.
SFT + RL (full) : 75.6 overall (+6.4 over SFT, improvements in all six categories).
Single-task specialists (six separate agents): 76.3 overall, slightly higher than the shared multi-task policy (75.6), but the shared policy covers all tasks with one model.
Reward ablations : removing video-quality reward drops to 73.3; removing tool-use reward drops to 71.2, confirming both signals are necessary.
Compositional Generalization: Procedural Identity (PI)
A zero-shot test combines procedural knowledge (PK) and single-subject identity (SI) – 100 unseen prompts pairing 20 professional actions with 20 specified subjects. No PI examples were seen during SFT, RL, or VABench. The RL-trained agent scores 73.2 , surpassing SFT (57.6) and nearing a hand-crafted fixed pipeline that explicitly chains text retrieval, image retrieval, and reference-conditioned generation (72.4). The RL agent invokes both text and image retrieval in 84% of trajectories vs. 40% for SFT, showing it learns to recombine tools for novel constraint combinations.
Extension to Robotics
An additional 1,000 robot manipulation examples (pick-and-place, towel folding) fine-tune the RL agent to call a motion predictor (π₀.₅) followed by a video generator (Ctrl-World). Qualitative videos show the policy sequentially invoking action prediction and video generation. This demonstrates the framework's extensibility to new tool categories, though quantitative success rates and real-robot execution remain unverified.
Conclusion
VideoGen-Agent demonstrates that training a multimodal agent to decide what information is missing and which tool can supply it substantially improves video generation on knowledge-intensive, identity-sensitive, physically grounded, and temporally structured tasks. The learned policy is reusable across generator upgrades and composes tools for unseen constraint combinations. Future work must address longer videos, more open-ended physics, and tool-call latency/cost trade-offs.
Paper: https://arxiv.org/abs/2609.24997 (arXiv:2609.24997)<br/>
Project page:
https://andyca111.github.io/VideoGen_Agent/Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
