JD Haibo's AI-Native Knowledge Base: 3-Layer Architecture, OKF Spec & 44% Faster Dev
JD Haibo built a three-layer AI knowledge base (project-domain matrix, role-specific views, executable skills) using the OKF specification, auto-generating structured docs from code, integrating with Harness for context-aware AI assistance, and validating via AB tests showing 44% faster requirement analysis with better coverage.
Background: The "Genius Intern" Problem
After integrating LLMs into daily development, the Haibo team found that models behave like a genius intern: smart but needing to rebuild system understanding from scratch every session. Three recurring pains emerged in their complex, multi-repo system:
AI cannot understand the system — searches millions of lines across dozens of repos, produces code that runs but violates architecture; call-chain breaks lead to guessing.
Impact analysis relies on tribal knowledge — only a few veterans know downstream effects of a field/interface change; they become bottlenecks and risks.
Test expectations are scattered — business rules, historical conventions, and hard-won lessons live in PRDs, chat logs, and personal memory.
The core problem: real system knowledge (architecture, rules, impact paths, gotchas) sits in heads and fragmented docs, unstructured and not consumable by AI. The goal: help the "genius intern" build rapid cognition by turning knowledge into searchable, traceable, inheritable, AI-consumable organizational assets.
Technical Selection: Prioritize Knowledge Production Over Retrieval
The team evaluated three retrieval paradigms:
Vector RAG — chunk docs, recall by semantic similarity; general, low integration cost.
GraphRAG — build knowledge graph, retrieve by relationships; excels at entity associations.
Agentic Search — model decides retrieval rounds; higher flexibility.
All three focus on "how to retrieve more accurately at runtime," but the real bottleneck was knowledge production . The team decided to shift "knowledge compilation" from runtime to maintenance time, adopting two industry-aligned approaches:
Karpathy's LLM Wiki — let models continuously curate knowledge into maintained Markdown.
Google Cloud's OKF (Open Knowledge Format) — a spec for packaging knowledge for AI agents.
OKF's three key rules:
One concept per Markdown file — metrics, tables, interfaces, flows, domain models atomically split; one file per concept with headings/tables/code blocks as semantic anchors.
YAML front-matter metadata — required type field lets AI identify purpose/ownership without reading full content.
System-reserved files — index.md for global navigation (progressive disclosure, avoids token explosion); log.md for change history and traceability.
OKF also supports lenient consumption : missing fields or broken links degrade gracefully to plain docs, enabling incremental migration of legacy materials.
Overall Architecture: Three-Layer Structured Knowledge
Based on OKF and their Harness automated delivery system, they designed a bottom-up, three-layer structure:
Layer 1: Project Knowledge × Domain Knowledge Matrix (Foundation)
Horizontal (Project Knowledge) : each code repo gets a {app}-knowledge-catalog (e.g., yzt-erp-wms-knowledge-catalog) with this structure:
yzt-erp-wms-knowledge-catalog/ index.md # root index · global navigation log.md # change/generation log flows/ # business flows, one doc per external interface chains/ # end-to-end call chains domain/ # domain models, DB entities & enums infrastructure/ # storage & middleware views/ # reverse index: which execution flows touch a resource (critical for impact analysis) references/ # external refs & cross-project dependencies 待澄清问题.md # AI-uncertain business points for experts to fillThe views/ (reverse index) is key for impact analysis: it maps every table, Redis key, MQ topic to its readers/writers and producers/consumers.
Vertical (Domain Knowledge) : a domain-knowledge/ tree stitches multi-repo interfaces into business-sequenced chains via MQ/JSF/scheduled tasks. First-level domains hold major chains; second-level domains hold detailed sub-chains. Domain knowledge references project details rather than duplicating them; when project knowledge changes, only the link layer updates — no conflicting copies.
Layer 2: Role-Specific Knowledge Views
The same foundation is re-organized for different roles without reinventing the base:
QA — business rules + test cases + historical pitfalls.
Product — requirement index, feature specs, business decisions.
Operations — operational rules, config strategies, SOPs.
The base matrix already contains the developer view (file:line-level implementation & call chains); the middle layer adds tailored views for non-dev roles.
Layer 3: Skill Capability Layer
High-frequency actions are packaged as Skills so AI can invoke a "project-aware action" instead of learning on the fly. Skills include:
generate-cross-project-deps (Dev) — Scan cross-repo JSF/JMQ dependencies, derive upstream/downstream relations, materialize domain knowledge into traceable business chains.
project-analysis-tool (Dev) — Analyze requirements against knowledge base: identify goals, scope, constraints; decompose into executable tasks with linked context.
qa-test-breakdown (QA) — Generate test cases (normal, exception, boundary) from PRD, business rules, historical cases; link to requirements & risks.
impact (Dev/QA) — Assess change impact via dependencies, call chains, data flows; classify d1/d2/d3 impact levels to guide regression & integration scope.
okf-knowledge-read (All) — Read & interpret knowledge base by intent; return summaries first, support progressive drill-down to docs, interfaces, chains, code.
Knowledge Initialization: Structuring Legacy Code
With thick legacy codebases, the team needed to bootstrap knowledge for full-lifecycle AI-assisted delivery. They split knowledge generation by endpoint and role because each stage needs different context: backend dev needs interfaces & call chains; frontend dev needs pages & interactions; architects need cross-system business panorama.
Backend Knowledge Generation: Extracting What Developers Actually Care About
When a backend engineer picks up code, they repeatedly ask: Entry (where does this flow start?), Chain (how does the call propagate?), Sink (which DB/middleware gets written?), Impact (who else is affected?). The generation pipeline extracts exactly these verifiable facts from code: build-knowledge-catalog — master skill: GitNexus index → enumerate all external interfaces → produce flows/ (one MD per interface) + services/domain etc. build-flow-chains — aggregate business chains on top of flows/ using three edge types: shared tables, MQ topics, task types. generate-cross-project-deps — cross-project JSF/JMQ dependencies → references/跨项目依赖.md (reuses project-analysis-tool).
Extraction rule: only write facts verifiable from code; leave blanks for unverifiable business semantics, flag them for human fill-in, then sample-verify. The resulting global call graph directly answers "who is affected by this change?" and becomes the foundation for impact analysis.
Example : a 400K-line repo ( yzt-erp-wms) produced ~1,051 docs: 1,001 flows/ (584 HTTP, 167 RPC, 140 MQ consumers, 96 jobs, 14 OpenAPI), 41 chains/, 3 ADRs, 1 domain model index (~59KB), 4 references. Each directory has an index.md; overall aggregation: interface details → business chains → architecture decisions → domain model.
Frontend Knowledge Generation: Reconstructing Page Operations & Calls
Frontend devs need: Operations (what can this page do?), Calls (which APIs, with what real parameters?), Permissions (who can see/click?), Collaboration (page navigation, micro-frontend split & communication). The build-frontend-knowledge-catalog skill auto-extracts these, organizing by page. Two hard parts: (1) real API input params — frontend often passes a single param object; fields must be inferred from every call site; (2) interactions & permissions — every button's action, validation, display condition must be exhaustive. Micro-frontend host/child split & communication is separately documented because it's the hardest for AI to reconstruct.
Example : middleground-react-web (micro-frontend host) yielded ~530 docs: 139 API modules (104 host, 35 child xcGood), 249 component docs (238 local, 7 hiboComponents, 4 private registry), 101 page docs (61 host, 40 child), 22 architecture/flow docs (15 business flows, 7 architecture overviews), 3 standards refs.
Cross-Project Knowledge: Stitch, Don't Rebuild
Domain knowledge consumes the project-level catalogs — "only stitch, never rebuild." It collects cross-project deps already produced, follows inter-system calls/messages, and assembles scattered chains into end-to-end business journeys, producing a cross-system business map.
Example : Aggregated Delivery domain knowledge base — 93 docs (10 indexes + 83 content): 25 core trading chains (dispatch 9, order mgmt 6, pricing 5, status sync 5), 51 base services & configs (base 10, dispatch config 11, dispatch search 9, wallet 10, capacity gateway 4, trade gateway 2, finance 2, open platform routing 1, biz adaptation 1, ops tools 1), 17 flows & overviews (12 end-to-end flows + 5 root overviews).
Harness Knowledge Engine: Embedding Knowledge Into the Delivery Pipeline
Knowledge base value lies not in standalone queries but in being woven into the entire R&D pipeline — continuously consumed and replenished.
Consumption Side: Precise Knowledge Lookup + On-Demand Code Indexing
Real workflow is longer than "requirement → coding"; multiple communication rounds precede coding:
业务定档 → 需求粗评 → 需求持续沟通确认 → 需求评审 → 研发规约生成 → 研发 TRD 生成 → 正式编码. At each step, different Skills consume the knowledge base: read knowledge to set direction + on-demand index code for details . Intermediate artifacts (meeting notes, diagrams, chats) are stored as informal knowledge in a temporary Harness project catalog as process assets for that requirement.
Concrete consumption via /okf-knowledge-read (6-step pipeline):
Step 0 Identify intent → business concept / service contract / data flow / arch decision? Step 1 Discover knowledge base → locate <cwd>-knowledge-catalog, read index + type-registry Step 2 Match concepts → grep coarse filter + semantic fine filter, lock entries Step 3 Load documents → summary first, full text if needed Step 4 Follow links → traverse associations (entity → service → data flow) Step 5 Group & inject → group by 7 Types, inject into contextDeposition Side: Two Artifacts → Distill → Human Confirm → Merge
After a requirement finishes, two knowledge sources exist: (1) master branch code (ground truth), (2) Harness process artifacts (specs, TRDs, communication logs). Both feed into /distill-catalog which distills into three knowledge dimensions:
Business/Project Knowledge — Project facts, clarified business rules, arch decisions, impact scope; domain terms, flow states, system boundaries, interface/data contracts, dependencies, decision rationale & trade-offs, change impact — for unified understanding & impact assessment.
Project Dev Experience — Pitfalls hit, conventions, reusable solutions; symptoms, root causes, fixes, avoid-list, team conventions, integration/deployment notes, reusable components/scripts/templates, best practices — to reduce repeated trial-and-error.
People-Centric Experience — Tech selection experience, review feedback, collaboration conventions; selection background, comparison dimensions, decision basis & applicability boundaries, review opinions & improvements, cross-team communication mechanisms, responsibility division, response SLAs, knowledge transfer methods — to improve collaboration & decision quality.
Human gate : distillation results are not auto-committed; humans review conflicts (owner decides), then merge into master knowledge branch. AI never overwrites autonomously — every update is auditable.
Bidirectional loop : validated process artifacts aren't throwaway; they consume knowledge base during generation, then get distilled back alongside code. Knowledge grows thicker and more accurate with each use.
Lifecycle: Anti-Rot Governance
Four-stage governance, moving toward platform automation:
① Create — Scan code/config/docs to generate baseline; every knowledge item bound to owner + Jing ME channel.
② Distill & Feedback — Post-requirement: master code (auto) + Harness artifacts (active) → distill-catalog → knowledge branches (project+domain+role) → human confirm → merge master knowledge branch (closed-loop engine).
③ Conflict Resolution — Distillation results human-confirmed; conflicts resolved by owner; confirmed merge to master knowledge branch.
④ Expiry Governance — Detect stale knowledge: periodic confidence-based review → renew valid, downgrade partially stale, archive fully stale.
Effect Validation: AB Experiment Proves Faster & More Comprehensive
Two sub-agents, identical prompts, same requirement, same repo, same methodology; only difference: Group A uses knowledge base, Group B uses raw code only. Task: cross-repo consumption chain + config whitelist-heavy requirement.
Requirement Analysis Time : Group A (Knowledge Base) 542 sec (~9 min) — built biz panorama via KB, then targeted 3 source files; short path, few wasted lookups. Group B (Code Only) 967 sec (~16 min) — no structured entry, manually inspected 16 source files; ~1.8× Group A.
Context Sources : Group A — 8 structured KB docs + 3 source files cross-verified; covered biz rules, arch decisions, interface chains, impact scope. Group B — Only 16 source files; info scattered in code, comments, config, naming; manual assembly of biz semantics & call relations.
Baseline Traceability : Group A — Every claim cites source (biz chain / line number); traceable, verifiable, auditable. Group B — Only "source code exploration"; lacks biz semantics & chain-level refs; re-reading code needed for review.
Result 2: More comprehensive (found hidden blockers). Group A, following cross-repo chains & config gates, uncovered two blockers invisible to single-repo code reading: (1) downstream consumer missing auto-retry; (2) production config whitelist silently fails. Across multiple requirements, knowledge base accelerated without sacrificing accuracy and used fewer tokens.
Testing Practice: Test Cases as Assets, Precise Regression Lookup
Solving three testing pains: cases scattered across platforms, disconnected from business, historical cases not reusable. Solution: test-knowledge/ under each domain holds role-specific knowledge + independent cases/ (functional / API / automation frameworks). Cases become reusable, evolvable, feedback-capable assets anchored to business knowledge.
Test Cases as Knowledge Assets: Business-Anchored, Bi-Directional, Composable
Business-anchored — cases belong to concrete domain/sub-flow; coverage check = one index lookup, no platform scraping.
Bi-directional flow — forward: /prd-diff-testcase generates cases from PRD + KB (reuse existing, update KB with new); reverse: /harvest-cases deduplicates & archives historical cases back to KB. One forward, one reverse — case library thickens with use, not chaos.
Composable automation — automation = atomic case + chain orchestration + variable threading . Each case is a minimal reusable execution unit (atom); chains orchestrate atoms into flows; variables thread through the whole chain — atoms reusable, chains composable.
Precise Testing: Cases Anchored, Diff-Driven Regression Scope
Cases carry business anchors + code anchors (interfaces, impl files/methods, touched tables). From a git diff, the KB traces affected cases:
git diff → changed files/interfaces/tables ├─ Direct hit: diff file ⊆ case anchor → MUST run ├─ Interface hit: changed interface → flows locate → reverse-find cases using it → SHOULD run ├─ Data hit: changed table/topic → views reverse-find interfaces → cases → SHOULD run └─ Chain spread: hit interface belongs to chain → pull entire chain's cases → RECOMMENDED runOutcomes: Precise (only truly relevant cases), No misses (reverse index + chain diffusion catches "change table A affects interface B"), Tiered (must/should/recommended gives regression priority).
Real-World Results: 4 Business Groups
Deployed in Haibo-Tian, Order Fulfillment, Product, Base — 4 groups. Testers ran real requirements through AI case generation, measuring Adoption Rate (adopted / AI-generated) and Coverage Rate (AI cases / final cases). Representative runs:
Base Trading — Meituan Mall Billing & Settlement — Adoption 95%, Coverage 100%
Order Fulfillment — Galaxy SOHO Auto Store Pushback — Adoption 100%, Coverage 95%
Haibo-Tian — Dada Delivery Alliance Integration · Finance — Adoption 90%, Coverage 95%
Haibo-Tian — Mixue Product Price Optimization — Adoption 100%, Coverage 91%
Order Fulfillment — Raw Material Supplier Contract Validation — Adoption 90%, Coverage 95%
Base Trading — POS Dish Estimation & Clearing — Adoption 90%, Coverage 90%
Adoption 90–100%, coverage up to 100%; testers shifted from "write from scratch" to "review & complete." The ceiling of AI-generated cases depends on whether the fed knowledge can understand the full business chain, and stays comprehensive and fresh.
Summary & Outlook
Circling back, the team solved the three opening knots: AI doesn't understand system, impact relies on memory, test expectations nowhere to find. They reframed "AI not working well" from "need a stronger model" to "how to manage context well." Artifacts produced:
One philosophy : Context Engineering — supply-side quality determines AI ceiling.
One architecture : AI-Native three-layer knowledge structure.
Two engines : GitNexus graph + dependency analysis engine.
One closed loop : Skill build/supplement/use + four-step lifecycle governance.
One controlled comparison : AB experiment proves with-KB is faster & more comprehensive.
Next evolution directions:
Platformization — from team tool to platform service, lowering adoption cost for more teams (in progress).
Automation — raise confidence of auto-generated contracts/chains, reduce manual distillation effort.
Knowledge finds people — flip from "people search knowledge" to "knowledge pushes to the right person at the right time."
The knowledge base is not the destination; it's the starting point for AI to truly "understand our system." When context supply is solid, every AI output aligns closer to the team's real engineering reality.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
JD Tech Talk
Official JD Tech public account delivering best practices and technology innovation.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
