Choosing Your Intelligent Data Querying Stack: Text-to-SQL vs Text-to-DSL vs Data Agent
The article compares three technical approaches for natural language data querying—Text-to-SQL, Text-to-DSL, and Data Agent—analyzing their principles, strengths, limitations, and enterprise selection criteria, while emphasizing semantic layers, multi-step analysis, and phased implementation.
Introduction
Traditionally, business users needed data analysts to write SQL and build reports for every metric. With large language models entering BI, intelligent data querying (智能问数) lets users ask questions in natural language and receive data, charts, and analysis conclusions directly. However, enterprise needs go beyond simple Q&A they require continuous exploration and reliable analysis workflows.
1. Text-to-SQL: Direct Natural Language to SQL
Text-to-SQL (NL2SQL) converts natural language into executable SQL. The basic flow: user question → understanding → SQL generation → execution → result return.
Example: "Who are the top 10 customers by sales in East China last month?" The system must identify time range, region filter, sales metric, and sorting, then generate SQL against the database schema.
Advantages: Simple technical chain, flexible development, quick natural language query implementation.
Two Underestimated Pitfalls
SQL Correctness ≠ Business Correctness
When a user asks "What was last month's sales revenue?" the system must decide among multiple ambiguous rules:
Order amount vs. actual paid amount?
Should cancelled orders be excluded?
How to deduct refunds?
Order time vs. payment time for attribution?
These rules cannot be inferred from field names alone.
Multi-Table Joins Cause Calculation Bias
If an order table joins multiple detail tables, improper handling duplicates amounts. Thus, Text-to-SQL's real difficulties lie in schema understanding, business definition recognition, and complex join handling.
In practice, enterprises also need to continue analyzing query results. FineBI Next illustrates this: users select existing data tables for AI analysis, which performs metric calculation, chart generation, key findings, and supports follow-up questions—turning a one-off query into continuous analysis.
2. Text-to-DSL: Introducing a Semantic Layer
Text-to-SQL requires the LLM to directly understand hundreds of tables and thousands of fields, which is unstable. Text-to-DSL adds a semantic layer: predefined metrics, dimensions, calculation rules, and data relationships. The LLM maps user questions to these business objects.
Typical flow: Natural language → business semantic parsing → DSL → query execution → result.
Example: Enterprise defines "Net Sales = Valid Order Paid Amount - Completed Refund Amount." Any future question about net sales reuses this rule, eliminating inconsistent calculations.
Three Key Advantages
Unified Definitions: Same metric reuses fixed calculation rules.
Controlled Queries: Restrict queryable business objects and operations.
Centralized Maintenance: Rule changes propagate to all related queries.
Costs and Trade-offs
Building the semantic layer requires upfront effort: metric systems, dimension models, business rules. Ongoing maintenance is needed when definitions change. A critical trade-off: stricter predefined rules reduce flexibility for ad-hoc exploration. If a user asks for a never-modeled metric, the system cannot answer directly. Therefore, semantic layer construction should prioritize high-frequency core metrics first.
3. Data Agent: Multi-Step Analysis Automation
Real business analysis often requires a sequence of tasks. Example: "Why did East China profit decline this month?" A single SQL only shows the magnitude. Explaining the cause needs:
Confirm Metric: Clarify profit definition, time range, comparison baseline, verify decline is real.
Decompose Profit Structure: Analyze revenue, cost, expense changes; calculate each factor's impact.
Drill Down Dimensions: Check region, product, customer, channel for anomaly concentrations.
Validate Hypotheses: Test if profit drop correlates with volume decline, price changes, product mix shift, or cost increases.
Form Conclusion: Summarize key drivers, charts, and follow-up verification items.
Data Agent addresses this with four capabilities:
Planning: Decompose complex questions into executable tasks.
Tool Calling: Invoke query, calculation, visualization tools per task.
Context Management: Retain prior analysis conditions and intermediate results.
Result Verification: Check execution results and adjust analysis path if needed.
Unlike single-turn querying, agents decide next steps based on intermediate results. FineBI Next's analysis agent demonstrates this: it breaks complex questions into continuous steps, generates charts, findings, and conclusions, and allows follow-up questions without re-selecting analysis direction. For structured tasks, a planning mode can pre-define steps before execution.
Higher autonomy demands stricter execution boundaries and result validation. An error in the first step (e.g., wrong metric) propagates through the entire chain. Thus, Data Agent competitiveness depends not just on operation count but on end-to-end reliability.
4. Comparing and Selecting the Right Approach
Enterprises can compare across five dimensions:
Comparison Across Five Dimensions
Core Capability : Text-to-SQL – natural language query; Text-to-DSL – semantic-governed query; Data Agent – multi-step analysis.
Build Cost : Text-to-SQL – relatively low; Text-to-DSL – high semantic modeling cost; Data Agent – complex system integration.
Query Flexibility : Text-to-SQL – high; Text-to-DSL – constrained by model scope; Data Agent – depends on tool capabilities.
Main Risk : Text-to-SQL – SQL and definition errors; Text-to-DSL – model maintenance lag; Data Agent – multi-step error accumulation.
Typical Scenario : Text-to-SQL – ad-hoc data retrieval; Text-to-DSL – standard metric Q&A Data Agent – business diagnosis.
These routes are not mutually exclusive. Data Agents can call Text-to-SQL tools and consume governed metrics from the semantic layer. Complex enterprises often need a combined architecture:
Standard operating metrics via semantic layer.
Ad-hoc exploratory questions via controlled SQL.
Complex business analysis delegated to Agent for planning and tool orchestration.
If a mature BI system exists, consider whether AI can leverage existing assets. FineBI Next enables business users to ask from an intelligent analysis entry, while analysts can invoke AI within existing projects for table creation, charting, dashboard adjustment, and content generation—embedding intelligent analysis into current workflows.
Three Self-Assessment Questions
Is the data foundation mature? Poor data quality or inconsistent core metric definitions undermine even the most advanced Agent.
What is the primary need? High-frequency fixed queries favor semantic governance; complex diagnostics need stronger planning.
Who owns ongoing maintenance? Any route requires dedicated ownership of metric rules, data permissions, test cases, and analysis processes.
Discussing technical sophistication without grounding in existing data foundations leads to mismatched investment and value.
5. Five Must-Solve Implementation Challenges
1. Unified Metric Definitions
Metrics like sales, profit, active customers need explicit formulas, statistical periods, data sources, and filter conditions. If departments define the same metric differently, AI cannot give consistent answers.
2. Inherit Existing Permission Systems
Intelligent querying accesses databases; must enforce row/column permissions, sensitive field protection, and audit trails—especially for finance and HR data.
3. Build a Standard Test Set
Curate 30–50 high-frequency business questions covering simple queries, complex joins, multi-turn follow-ups, and edge cases. Verify not only SQL execution success but also final result correctness.
4. Ensure Analysis Traceability
Retain key filter conditions, calculation bases, and analysis steps to avoid un-auditable conclusions. FineBI Next displays analysis steps, charts, key findings, and can generate dashboards or storyboards for review and sharing. Verified analysis methods can be saved as reusable skills for recurring tasks like sales reviews or profit analysis.
5. Human Confirmation for High-Risk Conclusions
AI-identified correlations (e.g., profit and cost) do not prove causation. Major business judgments require combining business events, external context, and other evidence. Intelligent querying shortens the analysis process but cannot replace necessary business judgment.
6. Recommended Three-Phase Rollout
Phase 1: Solve High-Frequency Queries
Pick one business domain (sales, finance, inventory), organize core metrics and standard questions, establish query accuracy baseline. Ensure common questions are answered reliably.
Phase 2: Extend to Complex Analysis
Introduce multi-turn follow-up, dimension drill-down, metric decomposition, chart generation. Focus on validating handling of context loss, definition drift, and error accumulation.
Phase 3: Institutionalize Reusable Analysis Capabilities
Solidify mature analysis flows into skills or business scenarios, gradually covering operations reviews, sales diagnostics, inventory analysis, etc.
Continuously monitor four metrics:
Result Accuracy: Does the final answer comply with business rules?
Task Completion Rate: Can complex analysis run through the full process?
Human Intervention Rate: Which steps still require manual handling?
Total Cost of Ownership: Model inference, data compute, and maintenance investment.
Only by tracking these can an enterprise judge whether intelligent querying delivers real business value.
Conclusion
From Text-to-SQL to Text-to-DSL to Data Agent, intelligent querying progressively expands its capability boundary. Text-to-SQL solves query flexibility; Text-to-DSL solves semantic consistency; Data Agent solves complex analysis execution. Enterprises should not blindly pursue the most complex architecture. Data scale, metric maturity, business needs, and maintenance capacity are the true determinants. Future competition will center on result accuracy, analysis traceability, and business scenario reusability—ultimately delivering an intelligent analysis system that understands business questions, uses data correctly, and continuously provides trustworthy insights.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Data Integration and Governance
Providing high-quality content on data integration and governance. Follow us!
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
