Doubao-Seed-2.1-pro Reads 232-Page Annual Reports: Building a Verified Financial Research Pipeline
The author tests Doubao-Seed-2.1-pro's multimodal and deep research capabilities by extracting structured financial data from 232-page annual reports, building a pipeline with cross-validation, Excel formulas, and HTML dashboards, then generalizing to BYD and Moutai, revealing that context understanding and evidence chains matter more than OCR accuracy.
Testing a Multimodal LLM on Real Financial Document Processing
On September 15, 2024, Doubao-Seed-2.1-pro was updated with enhanced multimodal understanding (40% faster image/video tasks) and deep research capabilities (evidence tracing, source authority verification, financial modeling). The author evaluated these claims by assigning a concrete, verifiable task: read a 232-page annual report from CATL (宁德时代) and produce a structured research workbook — extracting key figures, preserving evidence, cross-checking, performing financial calculations, and outputting both an Excel workbook and an HTML dashboard.
Pipeline Design: From PDF to Verified Workbook
The pipeline avoids feeding the entire PDF at once. Instead, it routes relevant pages, sends screenshots to the model for visual extraction, and demands strict JSON output. A sample JSON schema used for each field:
{
"key": "revenue",
"label_cn": "营业收入",
"period": "2025FY",
"value_thousand_cny": 423701834,
"unit": "千元",
"page": 11,
"confidence": 0.0,
"evidence": "截图中完整可见的行名、年份和金额"
}A hard rule was added: if the model cannot read a value clearly, it must return null — never infer from adjacent years. The pipeline then cross-checks every extracted field against the PDF text layer, validates 24 segment items, and writes three-statement articulation checks (assets = liabilities + equity, etc.) as Excel formulas for automatic verification.
First Run: CATL (232 Pages)
12 visual extraction calls were made (2 timeouts). Results: 70/70 financial fields matched the PDF text layer, 24/24 segment cells matched, 6/6 three-statement checks passed. Total tokens: 50,792; sequential latency: 10 minutes 24 seconds. The Excel workbook contains 7 sheets with formula-driven growth rates and ratios, re-verified with LibreOffice. The HTML dashboard consumes the same structured JSON.
Core metrics extracted:
Revenue: 4,237.02 billion (2025) vs 3,620.13 billion (2024), +17.04%
Net profit attributable to shareholders: 722.01B vs 507.45B, +42.28%
Operating cash flow: 1,332.20B vs 969.90B, +37.35%
Total assets: 9,748.28B vs 7,866.59B, +23.92%
However, the page coordinates were hard-coded for CATL — changing companies would break the routing.
Generalization: BYD (268 Pages) Reveals Hidden Assumptions
Wrapping the pipeline into a web assistant with automatic report fetching (via CNINFO) and anchor-based page routing (locate table headers, then follow continuation pages), the author tested BYD. The routing worked, but 27 "auto-corrections" appeared — all in the income statement and cash flow statement. Root cause: BYD declares the unit "千元" only once at the start of the financial statements section; subsequent pages lack unit markers. The original logic assumed each page repeats the unit, defaulting to "元" when missing, causing a 1000x discrepancy. Fix: determine unit per statement by scanning all pages of that statement.
Other BYD-specific differences:
Parent company balance sheet named "公司资产负债表" (no "母") vs CATL's "母公司资产负债表" → anchor set generalized to a header collection.
Line item naming variations: "投资活动产生的现金流量净额" vs "投资活动产生/(使用)的现金流量净额" → added label variants.
Period labels: "期初/期末" vs "年初/年末" or even "净减少额" → same handling.
Negative numbers: CATL uses minus sign (-649,350); BYD uses parentheses (649,350) → added parentheses parsing.
Decimal concatenation bug: BYD's key metrics page uses "元" with two decimals; the text-layer de-whitespace step merged adjacent numbers (e.g., 29,445,703,000.00 and 36,982,887,000.00 became 29,445,703,000.0036). Fix: restrict accounting decimals to 1-2 places — domain knowledge that accounting amounts never have three decimals.
After fixes: BYD 79/79 fields matched, 6/6 checks passed. Dashboard shows revenue 8,039.65B (+3.5%), net profit 326.19B (-19.0%), matching actual 2025 results.
Edge Case: Kweichow Moutai (143 Pages) Tests Continuation Logic
Moutai places the equity tail of the consolidated balance sheet and the parent balance sheet header on the same page. The continuation rule "stop if header appears in first 25% of page" incorrectly discarded the equity section (header at 224 characters), causing 2 of 6 articulation checks to fail. Rule corrected to "continue if real data rows exist before header". Thanks to the decision to persist every page's raw VLM response from the start, the 10 already-processed pages were replayed at zero token cost; only 2 new pages required fresh calls. After correction: Moutai 79/79 matched, 6/6 passed. Regression on CATL and BYD using stored responses confirmed no regressions.
Replay Mechanism: Turning AI Pipeline into Testable Software
Persisting raw VLM responses gave the system a "fixed test sample" akin to traditional software testing. Rule changes, parser fixes, or dashboard redesigns can be validated instantly by replaying historical responses without re-invoking the model. This replay mode became a first-class feature in the web assistant (toggle to use cached responses, zero tokens).
Three Key Findings
1. Multimodal Difficulty Is Context, Not OCR
A case study only serves one company; a tool must serve all. CATL succeeded on the first night, but BYD exposed "dialects" in page layout, unit declaration, label naming, negative formatting, and decimal handling within 10 minutes. Generalization value lies in surviving the second company's punch.
2. Deep Research Requires Evidence Chains, Not Just Search
Persisting raw responses per page was not for audit — it was for development efficiency. Every rule fix could be re-run at zero marginal cost. Without this, each iteration would burn full token budget, making experimentation prohibitively expensive.
3. Financial Systems Need Models That Can Say "I Don't Know"
The model should not replace analysts' buy/sell judgments (industry cycles, valuation assumptions, forward expectations). But a pipeline that returns null when uncertain, flags mismatches, and alerts on failed articulation checks provides a verifiable execution layer. In production, users need to know the provenance and confidence of every number. A system that dares to say "I don't know" is more valuable than one that always hallucinates an answer.
All tests used Volcano Engine's doubao-seed-2-1-pro-260915 via Responses API for visual extraction, on public annual reports only.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
AI Code to Success
Focused on hardcore practical AI technologies (OpenClaw, ClaudeCode, LLMs, etc.) and HarmonyOS development. No hype—just real-world tips, pitfall chronicles, and productivity tools. Follow to transform workflows with code.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
