JD Health Replaces RAG with Deterministic Agentic Search for Medical AI
JD Health rebuilt its medical AI stack across four layers, replacing probabilistic RAG with deterministic Agentic Search and structured knowledge base LLMwiki to achieve verifiable, auditable clinical task execution.
Problem: Probabilistic AI Fails in High-Stakes Medicine
Medical AI is shifting from answering questions to completing tasks, but in healthcare every output must be deterministically correct. General LLMs deployed directly suffer from uncontrollable probabilistic retrieval, a tendency to answer rather than execute, and generic engineering frameworks that cannot meet domain requirements — errors go undetected and unaccounted for.
Root Cause: Engineering Stack Determinism, Not Model Capability
The bottleneck is not model strength but the determinism of the entire engineering stack from model to knowledge to engineering to delivery. JD Health reframed the problem: achieving deterministic correctness requires rebuilding the full stack, not just improving the model.
Solution: Four-Layer Engineering Stack Reconstruction
Model Layer: Execution-Level Medical Code Model
Upgraded from a reasoning model to an execution-level medical code model. Code becomes the carrier of task execution, producing outputs that are verifiable and traceable.
Knowledge Layer: Deterministic Agentic Search & LLMwiki
Replaced probabilistic RAG with a self-developed deterministic Agentic Search and a structured knowledge base called LLMwiki. Knowledge is retrieved by following logical paths, making sources traceable and the verification process auditable.
Harness Layer: Directed Harness (MedWork)
Built a custom harness named MedWork that organizes tasks and deliverables, constrains tool calls, and validates critical medical logic.
Delivery Layer: Closed-Loop Evaluation & Recovery
Established a closed loop with evaluation, execution records, error recovery, and human review. Checkpoints and recovery mechanisms are set for long-running tasks; critical results undergo professional audit.
Components & Clinical Coverage
The stack is powered by four self-developed components: 京医千询 3.0, MedCode, LLMwiki, and MedWork. Together they support diverse clinical tasks including evidence retrieval, medical record summarization, outpatient assistance, recording-assisted diagnosis, AI-MDT, clinical nutrition, medication assistance, imaging, and research.
Evaluation Methodology: Three-Stage Assessment
The approach introduces a three-stage evaluation framework — model scores, task success, and actual adoption — to reduce judgment bias. It emphasizes building evaluation and regression checks around failure cases.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DataFunSummit
Official account of the DataFun community, dedicated to sharing big data and AI industry summit news and speaker talks, with regular downloadable resource packs.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
