JD Haibo's AI Knowledge Base: 3-Layer Architecture, OKF Spec & 44% Faster AB Test
JD Haibo's team built a structured AI knowledge base using OKF specification, a three-layer architecture (project-domain matrix, role-based views, skills), and a closed-loop lifecycle with automated generation and human validation; AB experiments showed 44% faster requirement analysis and more comprehensive coverage of cross-system dependencies.
Background: The "Genius Intern" Problem
After integrating LLMs into daily R&D workflows, the Haibo team at JD found that models act like a "genius intern" — smart but lacking persistent business context. Every session required rebuilding system understanding from scratch. Three recurring pains emerged:
AI cannot understand the system — searches millions of lines across dozens of repos, produces code that runs but violates architecture; call chains break and the model guesses.
Impact analysis relies on tribal knowledge — only a few veterans know downstream effects of a field/interface change; they become bottlenecks and risks.
Test expectations are scattered — business rules, historical conventions, and hard-learned lessons live in PRDs, chat logs, and personal memory.
The core problem: historical business knowledge is unstructured and not consumable by AI. The goal: turn knowledge into searchable, traceable, inheritable, and AI-consumable organizational assets .
Technical Selection: Prioritize Knowledge Production Over Retrieval
The team evaluated three retrieval paradigms:
Vector RAG — chunk documents, recall by semantic similarity; general-purpose, low integration cost.
GraphRAG — build knowledge graphs, retrieve by relationships; excels at expressing entity associations.
Agentic Search — model autonomously decides retrieval rounds; higher flexibility.
All three focus on "how to retrieve more accurately at runtime", but the real bottleneck is knowledge production : knowledge must be structured, bound with relationships, and precisely locatable. The decision: move "knowledge compilation" from runtime to maintenance phase.
Two industry references aligned with this view:
Karpathy's LLM Wiki — let models continuously maintain knowledge as Markdown.
Google Cloud's OKF (Open Knowledge Format) — a specification for packaging knowledge for AI agents.
Adopting OKF, the team organizes knowledge into interlinked Markdown trees at ingestion time; runtime only needs "navigation + extraction", avoiding on-the-fly chunking.
OKF Specification — Three Key Rules
One concept per Markdown file — metrics, data tables, interfaces, flows, domain models atomically split; one file per concept, using headings/tables/code blocks as semantic anchors.
YAML front-matter metadata — a YAML block at file top; type is the only required field, letting AI identify document purpose and ownership without reading full content.
System-reserved files — index.md as global table of contents for progressive disclosure (read index first, then fetch precisely, avoiding token explosion); log.md records changes for traceability.
OKF also supports tolerant consumption : missing fields or broken links don't error, gracefully degrading to plain documents, enabling incremental migration of legacy assets.
Overall Architecture: Three-Layer Knowledge Structure
Built on Harness automated delivery infrastructure, the knowledge base is not a flat doc dump but a structured, bottom-up growing three-layer architecture :
Bottom Layer: Project Knowledge × Domain Knowledge Matrix
The foundation, most objective and trustworthy, directly derived from code repos and call-chain analysis — a cross-cutting matrix.
Horizontal (Project Knowledge) : each code repo gets a {app}-knowledge-catalog (e.g., yzt-erp-wms-knowledge-catalog) capturing its own flows, views, domain models. Directory structure:
yzt-erp-wms-knowledge-catalog/ index.md # root index · global navigation log.md # change/generation log flows/ # business processes, one doc per external interface chains/ # end-to-end call chains domain/ # domain models, DB entities & enums infrastructure/ # storage & middleware views/ # reverse views: which execution flows access a resource (critical for impact analysis) references/ # external refs & cross-project dependencies 待澄清问题.md # business points AI unsure about, awaiting human fill-inThe views/ (reverse index) is key for impact analysis: readers/writers of a table, Redis key, or MQ topic are directly queryable.
Vertical (Domain Knowledge) : a business thread stitches interfaces from multiple project repos in temporal order, connected via MQ/JSF/scheduled tasks. First-level domains hold major chains; second-level domains hold detailed chains. Domain knowledge does not duplicate project details — only references them; when project knowledge changes, only the link layer updates, avoiding conflicts.
Middle Layer: Role-Based Knowledge Views
Same bottom matrix, reorganized by role without reinventing the base — one foundation, multiple perspectives:
Test : business rules, test cases, historical pitfalls.
Product : requirement index, feature specs, business decisions.
Operations : operational rules, config strategies, SOPs.
The bottom matrix already contains the R&D perspective (file:line-level implementation and call chains); the middle layer adds specialized views for non-R&D roles.
Top Layer: Skill Capabilities
High-frequency actions encapsulated as Skills, so AI doesn't relearn each time but invokes a "project-aware action". This layer is the core value-output interface. Skills include:
generate-cross-project-deps (Target: R&D) — Scan cross-repo JSF/JMQ dependencies, derive upstream/downstream relations, build domain knowledge, stitch into traceable business chains.
project-analysis-tool (Target: R&D) — Analyze requirements using knowledge base: identify goals, scope, constraints; decompose executable points; generate spec tasks with context.
qa-test-breakdown (Target: Test) — Break down test scenarios from PRD, business rules, historical cases; generate cases covering normal, exception, boundary; link to requirements & risks.
impact (Target: R&D/Test) — Assess change impact scope via dependencies, call chains, data flows; classify d1/d2/d3 impact levels; guide regression & integration scope.
okf-knowledge-read (Target: All) — Read & interpret knowledge base by user intent; return summaries & key conclusions first; support progressive drill-down to docs, interfaces, chains, code.
Knowledge Initialization: Structuring Legacy Code
During AI-Native transformation, teams face thick legacy codebases. To enable full-lifecycle automation via Harness, historical business knowledge must be initialized for every stage.
Why Split by Endpoint & Role?
AI can only replace a human in a stage if it understands that stage's context : backend coding needs interfaces & call chains; frontend needs pages & interactions; solution design needs cross-system business panorama.
Three knowledge libraries serve distinct consumers:
Backend — Serves: Backend R&D / AI Coding. Core problem solved: How interfaces are called, parameter/protocol contracts; call-chain flow through services/middleware; upstream callers, downstream dependents, data stores affected by a change.
Frontend — Serves: Frontend R&D / AI Coding. Core problem solved: Page operations, components, state; APIs called, real request/response payloads; role/condition-based visibility & degradation.
Domain (Business) — Serves: Architecture / Solution Design. Core problem solved: Cross-system business panorama, core domain objects & boundaries; end-to-end chain state transitions; impact of solution changes across systems, roles, scenarios.
Backend Knowledge Generation (400k LOC Repo Example)
Backend developers repeatedly need: Entry (where feature starts), Chain (call flow), Landing (which DB/middleware persists data), Impact (who is affected by a change). These answers are scattered across thousands of lines; generation extracts them automatically into queryable knowledge.
Extraction rule : only write facts verifiable from code; leave blanks rather than hallucinate; auto-extraction covers deterministic parts, ambiguous business semantics are flagged for human fill-in, then sampled for verification.
These facts form a global call-relation network, directly answering "who is affected by this change?" — the foundation for impact analysis.
Generated artifact (yzt-erp-wms) : ~1,051 docs across five categories:
flows/ Interfaces & Processes — 1,001 docs (core): HTTP 584, RPC 167, MQ consumers 140, JOB tasks 96, OpenAPI 14.
chains/ Business Chains — 41 docs covering procurement, allocation, inventory, replenishment, returns, supplier, damage, shelf-life.
decisions/ Architecture Decisions (ADR) — 3 docs: layered architecture, AOP aspects/interceptors, OpenApiType classification.
domain/ Domain Model — 1 doc (index.md, ~59KB summary).
references/ References — 4 docs: cross-project deps, type registry, JD platform tech stack.
Each directory has index.md for navigation. Overall hierarchy: "interface/process details → business chains → architecture decisions → domain model" forming a complete WMS knowledge graph.
Frontend Knowledge Generation (Micro-Frontend Example)
Frontend developers need: Operations (what actions, buttons, entry points), Calls (which APIs, real parameters), Permissions (role/condition visibility), Collaboration (page navigation, micro-frontend split & communication).
Extraction challenges: real API input params (frontend often passes a single param object; fields must be inferred from each call site), interactions & permissions (every button's action, validation, display condition), and micro-frontend split/communication (hardest for AI to infer alone).
Generated artifact (middleground-react-web) : ~530 docs (18 indexes + 512 content) across five categories:
API Modules/Interface Docs — 139 docs (core): main app 104, sub-app xcGood 35.
Component Docs/UI Components — 249 docs: local components 238 (main 151 + xcGood 87), hiboComponents 7, private registry 4.
Page Modules/Page Docs — 101 docs: main 61, xcGood 40.
Architecture Overview/Architecture & Flows — 22 docs: business flows 15 (end-to-end chains), architecture summaries 7 (main-sub relations, infra, inter-app comm, tech stack matrix, cross-repo deps).
Specs/References — 3 docs: UI spec, dev spec, component index.
Cross-Project Knowledge Generation: Stitch, Don't Rebuild
Solves the "single repo can't see end-to-end" problem. Domain knowledge derives from bottom frontend/backend outputs — "only stitch, don't rebuild": collect cross-project deps, follow inter-system calls/messages, connect scattered chains into end-to-end business journeys, producing a cross-system business panorama.
Example: Aggregated Delivery Domain — 93 docs (10 indexes + 83 content) in three categories:
Core Trading Chains — 25 docs: order placement 9, order mgmt 6, pricing 5, status sync 5.
Base Services & Config — 51 docs: base 10, delivery config 11, delivery search 9, wallet 10, capacity gateway 4, trade gateway 2, finance data 2, open platform routing 1, biz adaptation 1, ops tools 1.
Business Flows & Overview — 17 docs: flows 12 (merchant onboarding, trade onboarding, placement, cancellation, exception, status callback, completion compensation, financial settlement, system panorama), root overview 5 (README, glossary, gaps, outline).
Harness Knowledge Engine: Embedding Knowledge Into the R&D Pipeline
Knowledge base value lies not in standalone calls but in being woven into the entire R&D pipeline, continuously consuming and depositing .
Consumption Side: Precise Knowledge Location + On-Demand Code Indexing
Real process: Business Finalization → Rough Requirement Review → Continuous Clarification → Requirement Review → Spec Generation → TRD Generation → Formal Coding. At each step, different Skills consume knowledge: read knowledge base for direction + on-demand index code for details .
Intermediate artifacts (meeting notes, diagrams, dialogues) are stored as informal knowledge in Harness temporary project directories as process assets.
Concrete consumption via /okf-knowledge-read (6-step pipeline):
Step 0 Identify Intent → Business concept / service contract / data flow / architecture decision? Step 1 Discover Knowledge → Locate <cwd>-knowledge-catalog, read index + type-registry Step 2 Match Concepts → Grep coarse filter + semantic fine filter, lock entries Step 3 Load Documents → Summary first, full text if needed Step 4 Follow Links → Traverse associations (entity→service→data flow) Step 5 Group & Inject → Group by 7 Types, inject into contextDeposition Side: Two Artifacts → Distill → Human Confirm → Merge
After requirement completion, two knowledge sources emerge:
Master branch code (ultimate truth).
Harness process artifacts (specs, TRDs, multi-round clarifications).
Both fed into /distill-catalog together, distilled into knowledge base (project + domain + role knowledge):
① master code ─┐ ├─► distill → knowledge branch (project+domain+role) ─► human confirm ─► merge master knowledge branch ✅② harness process artifacts─┘Three knowledge dimensions distilled:
Business/Project Knowledge — Project knowledge, clarified business rules, architecture decisions, impact scope; domain terms, flow states, system boundaries, interface/data contracts, dependencies, decision rationale & trade-offs, change impact ranges — for unified understanding & impact assessment.
Project Dev Experience — Pitfalls hit, conventions, reusable solutions; symptom, root cause, resolution, avoidance checklist, team conventions, integration/deployment notes, reusable components/scripts/templates, best practices — to reduce repeated trial-and-error.
People-Centric Experience — Tech selection experience, review feedback, collaboration conventions; selection background, comparison dimensions, decision basis & applicability boundaries, review opinions & improvements, cross-team communication mechanisms, responsibility division, response SLAs, knowledge transfer methods — to improve collaboration efficiency & decision quality.
1️⃣ Human Confirmation : Distilled output not auto-committed; human reviews first, then merges to master knowledge branch. AI never overwrites autonomously; humans hold the final gate — every update is trustworthy & traceable.
2️⃣ Bidirectional Closed Loop : Verified process artifacts aren't throwaway; they consume knowledge base at generation, then distill back into the library alongside code. Knowledge is "fed while being used" — grows thicker and more accurate with each cycle.
Lifecycle: Knowledge Anti-Corrosion
To prevent staleness, four key processes (being platformized):
① Create — Scan code/config/docs to generate baseline; each knowledge item bound to owner + JD ME channel.
② Distill & Feedback — After requirement done, master code (auto) + harness artifacts (active) both distilled via distill-catalog back to library (closed-loop engine).
③ Conflict Confirmation — Distilled results human-confirmed; conflicts resolved by owner; confirmed merge to master knowledge branch.
④ Expiration Governance — Detect stale knowledge: periodic confidence-based renewal, partial downgrade, full expiration archival.
Effect Validation: AB Experiment Proves "Faster & More Comprehensive"
Two sub-agents with identical prompts, same requirement, same repo, same methodology ran in parallel for solution design; only difference: Group A used knowledge base, Group B used only code.
Result 1: ~44% Faster (cross-repo consumption chain + config gateway heavy requirement):
Requirement Analysis Time : Group A (Knowledge Base) 542 sec (~9 min): built business panorama via KB, then targeted verification of 3 source files; short path, few wasted lookups. Group B (Code Only) 967 sec (~16 min): no structured entry, manually inspected 16 source files; ~1.8× Group A.
Context Sources : Group A — 8 structured KB docs + 3 source files cross-verified; covered biz rules, arch decisions, interface chains, impact scope. Group B — Only 16 source files; info scattered in code, comments, config, naming; manual assembly of biz semantics & call relations.
Baseline Traceability : Group A — Every claim sourced (biz chain / line number); traceable, verifiable, auditable. Group B — Only "source code exploration"; lacks biz semantics & chain-level refs; re-reading code needed for review.
Result 2: More Comprehensive (Hidden Blockers Found) — Group A traced cross-repo chains & config gateways to discover two blockers invisible to single-repo code reading:
Cross-repo consumer missing auto-retry.
Production config whitelist gate fails silently.
Across multiple requirements, knowledge base improved speed without sacrificing accuracy, and reduced token consumption.
Testing Practice: Test Cases as Assets, Precise Regression Lookback
Solves three testing pains: cases scattered across platforms, misaligned with business, historical cases not reusable.
Solution : test-knowledge/ under each business domain holds role knowledge + independent cases/ layer (functional / API / automation framework). Cases no longer float in platforms; they live on business knowledge. Two bidirectional flows:
Anchored to Business : cases belong to concrete domain/sub-flow; coverage check = one index lookup, no platform crawling.
Bidirectional Flow : Forward /prd-diff-testcase generates cases from PRD+KB — reuse existing, update KB with new; Reverse /harvest-cases deduplicates & archives historical cases back to KB. One forward, one reverse — case library thickens with use, not chaos.
Composable Automation : Automation = case atoms + chain orchestration + variable threading . Each case is a reusable minimal execution unit (atom); orchestrated by business chains; variables thread through entire chain — atoms reusable, chains composable.
Precise Regression: Cases Anchored, Diff Traces Scope
Cases anchored to business & code anchors. Via relationships, trace to involved interfaces, implementation files/methods, touched tables. Incoming code diff triggers reverse lookup:
git diff → changed files/interfaces/tables├─ Direct hit: diff file ⊆ case anchors → Must Run├─ Interface hit: changed interface → flows locate → reverse query referencing cases → Should Run├─ Data hit: changed table/topic → views reverse query using interfaces → cases → Should Run└─ Chain diffusion: hit interface belongs to which chain → pull entire business chain cases → Suggest RunOutcomes: Precise (only truly relevant cases), No Leaks (reverse index + chain diffusion covers hidden impacts like "change table A affects interface B"), Tiered (Must/Should/Suggest — regression priority).
Real-World Adoption: 4 Business Groups
Deployed in Haibo Fields, Order Fulfillment, Product, Base — 4 groups. Testers ran real requirements through AI case generation, measuring:
Adoption Rate (adopted / AI-generated)
Coverage Rate (AI cases / final cases)
Results:
Base Trading — Meituan Mall Billing & Settlement: Adoption 95%, Coverage 100%
Order Fulfillment — Galaxy SOHO Auto Store Pushback: Adoption 100%, Coverage 95%
Haibo Fields — Dada Delivery Alliance Integration · Finance: Adoption 90%, Coverage 95%
Haibo Fields — Mixue Product Price Optimization: Adoption 100%, Coverage 91%
Order Fulfillment — Raw Material Supplier Contract Validation: Adoption 90%, Coverage 95%
Base Trading — POS Dish Estimation & Clearance: Adoption 90%, Coverage 90%
Adoption 90–100%, coverage up to 100%; testers shifted from "write from scratch" to "review & complete". The ceiling of AI-generated cases depends on whether fed knowledge understands full-chain business, is comprehensive, and stays fresh.
Summary & Outlook
Circling back, the team solved the three opening knots: AI not understanding system, impact analysis relying on memory, test expectations having no home. The problem "AI not working well" was reframed from "need a stronger model" to "how to manage context well". Along this path they accumulated:
One Philosophy : Context Engineering — supply-side quality determines AI ceiling.
One Architecture : AI-Native Three-Layer Knowledge Structure.
Two Engines : GitNexus Graph + Dependency Analysis Engine.
One Closed Loop : Skill Build/Supplement/Use + Four-Step Lifecycle Governance.
One Controlled Comparison : AB Experiment Proves Knowledge Base Delivers Faster & More Comprehensive Results.
Next evolution directions:
Platformization : Upgrade from team tool to platform service, lowering adoption cost for more teams (in progress).
Automation : Increase confidence of automated contracts/chains layer, reduce manual distillation cost.
Knowledge Finds People : Flip from "people search knowledge" to "knowledge proactively finds people", pushing relevant knowledge at the right moment.
The knowledge base is not the destination, but the starting point for AI to truly "understand our systems". When context supply is solid enough, every AI output aligns closer to the team's real engineering reality.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
JD Tech
Official JD technology sharing platform. All the cutting‑edge JD tech, innovative insights, and open‑source solutions you’re looking for, all in one place.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
