R&D Management 12 min read

AstraFlow Xingtu Client: Automating Systematic Literature Reviews with Evidence Gates

The article demonstrates how the AstraFlow Xingtu client automates systematic literature reviews by generating database-specific queries, retrieving and deduplicating records, applying evidence gates to prevent premature conclusions, and delivering structured outputs including evidence matrices and Zotero-ready citations, with human oversight at critical decision points.

UCloud Tech
UCloud Tech
UCloud Tech
AstraFlow Xingtu Client: Automating Systematic Literature Reviews with Evidence Gates

01 Input Research Question and Selection Criteria

The example research question asks whether SGLT2 inhibitors reduce heart failure hospitalization and cardiovascular mortality as of December 31, 2025, and whether effects are consistent across HFrEF, HFmrEF, and HFpEF subtypes. The input also specifies a time range, inclusion criteria (RCTs, systematic reviews, meta-analyses, formal guidelines), and exclusion criteria (animal studies, cell studies, diabetes-only studies, commentaries, press releases). Clear problem definition sets the boundaries for subsequent retrieval and evidence appraisal.

Research protocol input: topic, time range, inclusion/exclusion criteria
Research protocol input: topic, time range, inclusion/exclusion criteria

02 Generate Query Plans and Execute Retrieval

The workflow generates database-specific query syntax rather than submitting identical keywords to all sources. OpenAlex uses its filter syntax, Europe PMC adds a FIRST_PDATE constraint, and Crossref uses its own query parameters. The query plan preserves each database's retrieval rules for later audit of inclusion and exclusion rationale.

Query planning: separate search strategies for three databases
Query planning: separate search strategies for three databases

03 Parallel Retrieval and Exception Logging

Retrieval tasks run in parallel across the three databases. In this run, Crossref failed and the failure status was recorded rather than skipped. OpenAlex returned results but the date range still required manual confirmation, so it was not marked as fully compliant. The system retains actual execution states to avoid recording exceptions as successes.

Three-database retrieval results: hit counts, status, and exception records
Three-database retrieval results: hit counts, status, and exception records

04 Distinguish Bibliographic Records from Direct Evidence

After initial screening, 12 candidate records were returned; deduplication left 11 unique records (1 duplicate removed). Records were classified by study type and identifier:

Formal guidelines: 2 (2021 ESC Heart Failure Guideline and 2023 ESC Guideline Update)

Systematic reviews/meta-analyses: 4

RCT primary results: 0 (only 1 trial design paper)

Direct matches to the research question: 2; verifiable direct outcome records: 0

The workflow distinguishes "bibliographic records" (confirmed by DOI, PMID, or guideline citation) from "direct evidence" that actually answers the specific research question.

Retrieval audit and evidence package: hit results, time range, and study type verification
Retrieval audit and evidence package: hit results, time range, and study type verification

05 Q1 Gate Not Passed, Conclusion Generation Paused

An evidence gate requires at least 3 verifiable direct evidence records before generating a definitive conclusion. This run scored 0.42 and failed the gate. Specific gaps:

Insufficient directly matched and verifiable records (need 3)

Missing effect sizes for heart failure hospitalization and cardiovascular mortality

Missing subgroup data for HFrEF, HFmrEF, HFpEF

Missing third-source verification

The gate prevents generating conclusions based solely on indirect guideline information when direct evidence is lacking.

Q1 gate decision: score 0.42, not passed; gaps in direct evidence, key effect sizes, and third-source verification
Q1 gate decision: score 0.42, not passed; gaps in direct evidence, key effect sizes, and third-source verification

06 Supplementary Retrieval and Human Confirmation When Evidence Is Insufficient

After gate failure, the workflow produces a targeted supplementary retrieval plan focusing on key trials such as DAPA-HF, EMPEROR, and DELIVER. Whether the supplementary results are merged into the current evidence package and processing continues requires human approval. The system can generate the next retrieval plan based on missing items, but execution and integration are human decisions.

Human confirmation during correction: whether to continue review generation is decided by a person
Human confirmation during correction: whether to continue review generation is decided by a person

07 Output Evidence Grading and Review Conclusion

The evidence matrix for this run is graded as follows:

C1: Heart failure hospitalization reduced → unsupported (insufficient evidence)

C2: Cardiovascular mortality reduced → unsupported

C3: Consistent effect across three heart failure types → unsupported

C5: Bibliographic identity of guidelines and trial designs → partial (metadata only confirmed)

The final conclusion states: "Evidence insufficient/pending verification, cannot be used for clinical decisions." The report lists next steps: obtain results from the four major RCTs, extract hazard ratios/relative risks with 95% confidence intervals, and verify subgroup interactions. Although no definitive answer was produced, the evidence gaps and next retrieval requirements are explicitly recorded for future supplementation.

08 Deliverables, Workflow Architecture, and Applicability Boundaries

On completion, the workflow outputs a complete JSON result containing search audit, evidence matrix, consensus and disputes, update log, and monitoring plan. It also generates 7 draft citations in RIS/CSL-JSON format as a Zotero handoff package. The package does not write directly to a Zotero library; it only prepares content for import, leaving the decision to import and verify to the user.

Deliverables: search audit, report results, and Zotero handoff package
Deliverables: search audit, report results, and Zotero handoff package

The workflow follows a defined SOP: convert the research question into a query plan, access multiple data sources in parallel, then deduplicate, verify identifiers, and package evidence. If evidence falls short of the gate, a supplementary retrieval plan is generated and proceeds only after human approval.

The architecture diagram shows inputs (research topic, time range, inclusion/exclusion criteria), processing nodes (query planning agent, OpenAlex/Europe PMC/Crossref retrieval branches, deduplication merge, evidence packaging, Q1 gate, supplementary retrieval, review generation), and outputs (report and Zotero handoff package).

Research information collection workflow architecture: from research question to report delivery
Research information collection workflow architecture: from research question to report delivery

Compared with typical manual processes, this workflow differs in three main ways:

Reduces repetitive organization: database switching, query syntax adaptation, initial deduplication, and citation formatting are automated; humans focus on problem definition, full-text verification, and result interpretation.

Encodes evidence requirements as process nodes: the workflow explicitly defines the number of direct evidence records required, mandatory fields to verify, and the supplementary retrieval method when evidence is insufficient.

Retains update rationale: search strings, hit counts, failure reasons, exclusion records, and update logs are preserved for handoff and future updates.

The Xingtu client supports combining models, agents, external APIs, and result files within the same workflow. When tasks change, individual nodes can be adjusted without rebuilding the entire environment.

This workflow suits scenarios requiring continuous literature retrieval, evidence organization, and citation maintenance. Its primary role is to standardize repetitive tasks such as searching, deduplication, evidence packaging, and citation handoff while preserving auditable and resumable process records. Judgments on evidence sufficiency, conclusion generation, merging of supplementary retrieval results, and whether results inform clinical or research decisions all remain human responsibilities.
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

systematic reviewevidence gateAstraFlow Xingtuevidence synthesisresearch workflow automationZotero integration
UCloud Tech
Written by

UCloud Tech

UCloud is a leading neutral cloud provider in China, developing its own IaaS, PaaS, AI service platform, and big data exchange platform, and delivering comprehensive industry solutions for public, private, hybrid, and dedicated clouds.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.