VisInteract: Recovering User Intent in Text-to-Vis via Interactive MCTS Search
VisInteract introduces a dynamic interactive Text-to-Vis paradigm with the VisInteract-Bench benchmark and Vis-MCTS method, using Monte Carlo Tree Search with progressive widening, cross-path memory, and dimension-aware reward decomposition to recover user intent from imperfect queries, achieving 51–57% Merge Task Score on 1,098 samples across 11 databases.
Research Background
Text-to-Visualization (Text-to-Vis) translates natural language queries into database queries and visual charts. While large language models are powerful, users rarely express their full intent in a single query. For example, a request like "plot monthly revenue of 2025 top-5 products by region" leaves ambiguities: top-5 by sales or volume? Are refunds included? Does the database even contain 2025 data? Existing systems typically generate a chart in one shot, silently choosing one interpretation.
Real-world enterprise scenarios demand multi-turn interaction to fill gaps, resolve ambiguities, and correct errors — effectively an intent-recovery process rather than single-shot generation. However, current Text-to-Vis research assumes queries are already complete, and evaluation benchmarks do not reflect imperfect inputs or open-ended visual outputs.
Key Challenges
Challenge 1: Benchmarks Assume Perfect Queries and Struggle with Open Outputs
Existing benchmarks (NVBench, VisEval, MultiVis-Bench) assume the user's intent is fully specified. They are mostly single-turn or provide only static, pre-written feedback. In contrast, real queries often contain ambiguity, missing information, or factual errors. Moreover, Text-to-Vis outputs are inherently open: the same analytical intent can be satisfied by different chart types, color mappings, or encodings. Exact-match evaluation penalizes valid but differently styled visualizations, conflating requirement satisfaction with aesthetic preference.
Challenge 2: Methods Cannot Actively Recover Intent Through Dynamic Interaction
Pipeline methods generate once and stop. Ambiguity-aware methods enumerate a fixed candidate set. Feedback-based variants wait for the user to initiate corrections and follow a single modification trajectory. None can backtrack from early wrong assumptions, systematically explore the design space, or converge reliably to the user's true intent.
Benchmark Construction: VisInteract-Bench
To address Challenge 1, the authors release VisInteract-Bench , the first benchmark for dynamic interactive Text-to-Vis under imperfect queries. It contains:
1,098 samples across 11 databases and 75 tables
7 domains: sports, entertainment, education, finance, healthcare, etc.
Average 4.7 key features per sample
~61% samples with ambiguity, ~69% with factual errors
Multi-turn, multi-modal interaction that adapts to system output (not pre-scripted)
A User Agent simulates the user: it knows the ground-truth key features, reference code, and reference chart but refuses direct intent disclosure. Text questions are answered only for the asked features; visual feedback only indicates what is wrong, not how to fix it. Text queries resolve data-layer ambiguities (filters, aggregations); visual feedback catches rendering issues that code inspection misses.
Evaluation uses a dual-perspective LLM-as-Judge approach: a code judge verifies whether generated code satisfies each key feature; a chart judge verifies whether the rendered chart satisfies each key feature. A feature passes only if both judges agree. The primary metric is Task Score : a sample succeeds only when all its key features are satisfied. This design evaluates requirement fulfillment, not pixel-perfect match to a reference chart.
Core Method: Vis-MCTS
To address Challenge 2, VisInteract models dynamic interactive Text-to-Vis as a partially observable sequential decision process. The system sees the imperfect query and database schema but not the true user intent, gaining indirect evidence through text questions and visual feedback. Single-chain ReAct cannot recover from early mistakes. The authors adopt Monte Carlo Tree Search (MCTS) and propose Vis-MCTS with three innovations addressing three "walls" of naive MCTS:
Wall 1: Infinite Legal Moves
At each step, the system could emit arbitrary SQL, chart code (Altair), or a natural language question. Expanding all children would explode the tree immediately.
Wall 2: Repeated Questions Across Paths
Classic MCTS treats each path independently; the same clarification (e.g., "sort by sales") would be asked repeatedly on different branches.
Wall 3: Coupled Total Score
User feedback like "data is correct, chart type is wrong" yields a single score (e.g., 2/10). Uniform back-propagation would penalize the correct SQL step for the chart error.
Component 1: Selection & Expansion with Progressive Widening
The search branches among four action types: query SQL, generate Altair chart, ask user, or terminate. A node expands new children only after being visited a certain number of times (progressive widening), with a hard cap. Diversity-promoting prompts and deduplication prevent the model from generating near-duplicate code in the same context.
Component 2: Simulation with Cross-Path Memory
Clarifications obtained and visual critiques received on one branch are written into two global memories: all text Q&A pairs, and all visual scores with critiques. Once a clarification is resolved on any path, the entire tree can reuse it, avoiding redundant questioning cost.
Component 3: Dimension-Aware Reward Decomposition
The user's holistic satisfaction score is decomposed into three dimensions: data fidelity, visual design correctness, and intent alignment. Each dimension's score is back-propagated only to the relevant node types (SQL nodes, chart nodes, question nodes). The decomposed scores are anchored to the original total score to prevent judge bias drift.
Example: User feedback "I want a bar chart sorted by area descending" with satisfaction 2/10. Decomposition yields data fidelity 10, visual design 3, intent alignment 2. The SQL step receives high reward; the chart-type step is penalized alone. Uniform back-propagation would have dragged down all three steps.
Experimental Results
Evaluated on VisInteract-Bench with two backbone LLMs: Qwen3.5-flash and Gemini-3.1-flash-lite-preview. Baselines include non-interactive (Self-Correction LLM, nvAgent) and interactive methods (ReAct, Best-of-N ReAct, MultiVis-Agent). All share aligned tools and the User Agent; only the search strategy differs.
Merge Task Score (both code and chart features satisfied): Vis-MCTS achieves 51.46% (Qwen) and 57.19% (Gemini), outperforming Best-of-N by +13.40% / +16.27% and non-interactive self-correction by >5× (7.38% / 10.66%).
Code Task Score : improves from best baseline 50.00% / 54.79% to 61.66% / 67.21%.
Chart renderability: 100% for all methods.
With a fixed budget of 10 rollouts, Best-of-N (10 independent chains) shows unstable gains: ~9% over single ReAct on Qwen but only ~1.3% on Gemini. Vis-MCTS benefits from prefix reuse, backtracking, and shared interaction memory.
Ablation Study (Qwen, Merge Task Score)
Full Vis-MCTS: 51.46%
Without visual feedback: 24.24% (largest drop — loss of terminal signal)
Without cross-path memory: 34.06%
Without reward decomposition: 43.26%
Without text questioning: 44.08%
The results confirm that the improvement stems from the synergy of branching/backtracking, interaction sharing, and error-attributed credit assignment, not any single trick.
Judge Consistency Verification
To ensure evaluation reliability, the authors cross-validated with a GPT-family judge and human experts. Pairwise Cohen's Kappa and three-way Fleiss' Kappa all exceed 0.81 ("almost perfect agreement"). Qwen-judge alignment with humans matches GPT-judge alignment, confirming that reported gaps are not artifacts of a single judge's preference.
Case Study
An illustrative interaction shows Vis-MCTS handling a query with both factual error (requesting 2025 data that does not exist) and chart-type mismatch (user wants bar chart sorted by area, not pie chart). Through successive text questions and visual feedback, the system converges to the correct visualization. The dimension-aware reward decomposition selectively penalizes only the chart-type step, preserving credit for correct data retrieval.
Conclusion & Outlook
VisInteract reframes Text-to-Vis as intent recovery from imperfect queries via dynamic interaction. VisInteract-Bench enables evaluation with ambiguous, incomplete, or erroneous inputs and open-ended visual outputs. Vis-MCTS provides a principled search mechanism that decides when to query, what to ask, and whether to revise data, design, or intent understanding after negative feedback. For enterprise users with data already in databases, multi-turn interaction to lock down the right view is more practical than one-shot generation.
Limitations: the User Agent is more cooperative than real users, and multi-turn search is more expensive. Future work includes incorporating more realistic user feedback, richer visualization grammars, and better explanation of "why this chart".
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
