JD Health Replaces RAG with Deterministic Agentic Search for Medical AI

JD Health rebuilt its medical AI stack across four layers, replacing probabilistic RAG with deterministic Agentic Search and structured knowledge base LLMwiki to achieve verifiable, auditable clinical task execution.

DataFunSummit
DataFunSummit
DataFunSummit
JD Health Replaces RAG with Deterministic Agentic Search for Medical AI

Problem: Probabilistic AI Fails in High-Stakes Medicine

Medical AI is shifting from answering questions to completing tasks, but in healthcare every output must be deterministically correct. General LLMs deployed directly suffer from uncontrollable probabilistic retrieval, a tendency to answer rather than execute, and generic engineering frameworks that cannot meet domain requirements — errors go undetected and unaccounted for.

Root Cause: Engineering Stack Determinism, Not Model Capability

The bottleneck is not model strength but the determinism of the entire engineering stack from model to knowledge to engineering to delivery. JD Health reframed the problem: achieving deterministic correctness requires rebuilding the full stack, not just improving the model.

Solution: Four-Layer Engineering Stack Reconstruction

Model Layer: Execution-Level Medical Code Model

Upgraded from a reasoning model to an execution-level medical code model. Code becomes the carrier of task execution, producing outputs that are verifiable and traceable.

Knowledge Layer: Deterministic Agentic Search & LLMwiki

Replaced probabilistic RAG with a self-developed deterministic Agentic Search and a structured knowledge base called LLMwiki. Knowledge is retrieved by following logical paths, making sources traceable and the verification process auditable.

Harness Layer: Directed Harness (MedWork)

Built a custom harness named MedWork that organizes tasks and deliverables, constrains tool calls, and validates critical medical logic.

Delivery Layer: Closed-Loop Evaluation & Recovery

Established a closed loop with evaluation, execution records, error recovery, and human review. Checkpoints and recovery mechanisms are set for long-running tasks; critical results undergo professional audit.

Components & Clinical Coverage

The stack is powered by four self-developed components: 京医千询 3.0, MedCode, LLMwiki, and MedWork. Together they support diverse clinical tasks including evidence retrieval, medical record summarization, outpatient assistance, recording-assisted diagnosis, AI-MDT, clinical nutrition, medication assistance, imaging, and research.

Evaluation Methodology: Three-Stage Assessment

The approach introduces a three-stage evaluation framework — model scores, task success, and actual adoption — to reduce judgment bias. It emphasizes building evaluation and regression checks around failure cases.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Medical AIJD HealthAgentic SearchRAG AlternativeClinical TasksDeterministic AIEngineering StackLLMwikiMedWork
DataFunSummit
Written by

DataFunSummit

Official account of the DataFun community, dedicated to sharing big data and AI industry summit news and speaker talks, with regular downloadable resource packs.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.