OmniVChat: Teaching Models Native Video Calls via Synthetic Data Generation
OmniVChat introduces a synthetic data pipeline (OmniVChat-Studio), benchmark (OmniVChat-Bench), and RL reward design (OmniVChat-RL) to train multimodal models for native audio-visual dialogue, achieving strong generalization from synthetic to real human interactions.
OmniVChat defines native audio-visual dialogue as models directly receiving user audio and video simultaneously — without external text questions, subtitles, or ASR — preserving cues like tone, pauses, facial expressions, and camera orientation. For example, a user holding a phone while asking "Is the meeting-room door on my left open?" embeds the question in both speech and visual context.
Problem: Missing Data and Evaluation Standards
Real user–model video-call recordings are scarce, especially multi-turn dialogues that include the model's own previous responses. Existing omni models therefore rely on auxiliary modules (ASR, etc.), adding compute and losing information. Moreover, a good reply must jointly consider environment, objects, and emotion; keyword matching cannot reliably judge quality.
Approach: Generation for Comprehension
The authors adopt a generation-for-comprehension strategy: first define target capabilities, then synthesize matching dialogues, and derive reference answers and scoring rubrics from the actual synthesized audio-visual content. The same rubric serves both evaluation and RL reward.
OmniVChat-Studio: Multi-Agent Data Engine
OmniVChat-Studio is a multi-agent system that takes a sub-category configuration and text corpora (as random seeds) and outputs single- and multi-turn audio-visual dialogues with reference replies and tiered rubrics. Four agents collaborate:
Director : conceives scenes, writes scripts/prompts, makes accept/reject decisions, produces reference replies and rubrics.
Renderer : renders approved prompts and reference media into audio-synchronized clips.
Reviewer : watches the rendered audio-visual output, writes content descriptions and quality reports, answers targeted verification questions.
Validator : deterministic rule checker for structured scripts/dialogue logs against configuration requirements.
Single-Turn Synthesis
Sampling starts by drawing language, scene, camera mode, audio-visual distractors, and three text corpora. The Director creates a story, refines the user's question, target object appearance, speech timing, and wait periods into a timeline-structured script. Validator checks format, timeline, speech rate, and config rules; Director revises. After approval, Director specifies shooting, continuous capture, and sync requirements for Renderer. Reviewer then checks actual speech, visual content, cuts, body language; Director judges usability, launching targeted Q&A for ambiguous details (e.g., door open/closed). Reference reply and rubric are created after acceptance, grounded in the actual audio-visual evidence.
Multi-Turn Synthesis
The system plans the full dialogue and media-reference relations upfront, then writes prompts turn-by-turn, reusing previously approved video, audio, or last frames. Earlier objects, speaker voices, and scene states become conditioning for later turns; location or topic changes can select appropriate historical references. Each turn's clip and reply enter history; after all turns, the full record is validated, with repairs starting from affected turns.
Extensibility
New sub-categories can add configuration, specific prompts, and rules via Flexible modules while reusing Fixed modules and the overall review pipeline. Output instances contain audio-visual data, reference replies, and rubrics for downstream evaluation and training.
OmniVChat-Bench: Scoring Native Audio-Visual Dialogue
The benchmark contains 2,800 synthetic dialogues (2,550 single-turn, 250 multi-turn) across 5 capability categories, 17 sub-categories, 22 scene domains :
Dialogue State & Link Perception (DSLP) : connection health, user turn completion, camera orientation; wait during thinking pauses, use actual user/camera position for left/right queries.
Multimodal Entity Alignment (MEA) : link spoken references to visual objects, identify the active speaker among multiple talkers, track previously seen objects and speaker changes across turns.
Model Self-Awareness (MSA) : correctly state identity and physical limits (e.g., no limbs → cannot move objects, offer executable alternatives).
Anti-Hallucination (AH) : replies must stay faithful to audio-visual evidence; note missing objects, correct user's factual errors (e.g., wrong color), avoid continuing on false premises.
Emotion Recognition (ER) : combine facial expression and voice to recognize emotion; same words spoken anxiously vs. calmly may warrant different responses.
Each sample includes a reference reply and a tiered rubric . Standards are organized hierarchically: a criterion scores only if it is hit and all lower-level criteria are hit. An LLM judge evaluates each criterion semantically; deterministic aggregation yields the final score. This allows phrasing variation while pinpointing exactly what a reply missed. For multi-turn evaluation, all models receive identical history clips and reference replies; only the final turn is scored, enabling fair comparison under the same context.
OmniVChat-RL: Turning Evaluation Rubric into Reward
OmniVChat-RL jointly optimizes correctness, efficiency, and style. The total reward for a sampled reply y is: R(y) = r(y) + 0.5*f(y) + 0.1*e(y) + 0.5*s(y) r(y) : tiered rubric score (correctness).
f(y) : format check — non-empty, no thinking tags.
e(y) : efficiency — rubric score divided by reply length, min-max normalized within the candidate group for the same question; shorter correct replies score higher.
s(y) : style — seven criteria from the training system prompt (natural spoken language, conciseness, role-play vs. fact distinction, grammar fluency, etc.); all seven must pass for score 1, else 0.
Training data OmniVChat-Bench-Train = 5,600 synthetic dialogues + 560 dev dialogues (all from Studio). Dev set selects checkpoints; 2,800 OmniVChat-Bench samples (non-overlapping with train) test training gains; OmniVChat-Bench-Human (real human recordings) tests real-world generalization only, never used in training or selection.
Base model: Qwen3-Omni-30B-A3B-Instruct Thinker (directly encodes video/audio, outputs text). Training: GSPO, rank-64 LoRA , frozen audio/visual encoders, 1,000 steps, checkpoint every 20 steps. Final checkpoint chosen at step 940 by dev-set Subcategory Mean; same checkpoint evaluated on both synthetic Bench and human Bench.
Results: Synthetic Training Generalizes to Real Interactions
Key finding: RL on purely synthetic dialogues substantially improves real-interaction quality. The selected checkpoint raises OmniVChat-Bench score from 0.465 to 0.652 and OmniVChat-Bench-Human from 0.402 to 0.632 (Δ = +0.230). Gains transfer beyond the synthetic distribution to handheld-device, natural-environment, real-speech inputs.
OmniVChat demonstrates a scalable recipe: generate to cover scarce scenarios, review to ensure evidence reliability, unify evaluation and reward via tiered rubrics, and independently verify with real recordings. Synthetic data does not replace real data — real data remains crucial for testing generalization, discovering missing capabilities, and calibrating generation distributions — but well-designed synthetic data can effectively train models for real-world native audio-visual dialogue when human interaction data is expensive or unavailable, providing a solid foundation for future integration of real feedback and complex multi-turn scenarios.
Resources
Paper (arXiv): https://arxiv.org/abs/2609.21465 Hugging Face Dataset: https://huggingface.co/datasets/Harland/OmniVChat GitHub Code:
https://github.com/HarlandZZC/OmniVChatSigned-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
