Does AI-Generated Documentation Count as Knowledge Accumulation? A Practical Framework
The article evaluates whether using coding agents to write docs, Jira tasks, MR descriptions, and Skills constitutes real knowledge accumulation, proposing four criteria — findable, trustworthy, usable, updatable — and a workflow linking task evidence, ADRs, runbooks, tests, and Skills with continuous validation.
The author noticed a job requirement: "when using coding agents, must be able to accumulate knowledge." They compared their own practices — having AI write documentation, Jira tasks, Git commit messages, and GitLab MR descriptions, plus creating Skills, adopting MCPs, and maintaining project context and Claude rules — and asked whether these count as knowledge accumulation.
Judgment: These practices already contain knowledge accumulation. What's missing are source traceability, applicability conditions, reuse entry points, and feedback loops. Volume of docs, Skills, or tools alone doesn't prove effectiveness; the real test is whether a future agent or teammate can find the knowledge, apply it correctly, and know when not to copy it.
1. What Happens After "It's Written Down"
Scenario: An agent fixes a release script and writes a detailed MR. A month later, a colleague modifies a neighboring script and hits the same issue. They don't know the old MR number or why a simpler approach was rejected. Even a 2,000-word MR fails to deliver value.
Contrast with a short module note:
This module retries external writes. Before modifying retry logic, read docs/decisions/007-write-retry.md and check the associated regression test. This decision only applies to interfaces that already provide idempotency guarantees.
This note connects three things: current modification location, past decision, and executable verification.
Four questions to judge effective accumulation:
Findable? Can relevant records be located via keywords, module paths, or task entry points?
Trustworthy? Are conclusions backed by code, tests, logs, or confirmed decisions?
Usable? Are preconditions, steps, stop conditions, and success criteria clearly written?
Updatable? When the environment changes, is someone responsible for revision or deprecation marking?
These are the author's working criteria, not an industry standard.
2. What Each Practice Actually Accumulates
Docs, Jira, Commits, MRs: Preserving Task Context & Evidence
Jira retains requirements, acceptance criteria, progress; commits link code changes; MRs capture implementation rationale, review discussions, verification results. AI lowers the cost of writing down what used to stay in heads.
But rich descriptions need verification. Agents may write "suggested tests" as "tests passed," speculation as root cause, or invent design rationales never discussed. The author asks AI to distinguish: observed phenomena, unverified hypotheses, executed checks, skipped checks, pending decisions. "Why this choice" must rely on actual discussion and evidence, not AI post-hoc storytelling.
MRs can hold long-term knowledge. If a decision affects future development, give it a stable entry (module doc or ADR) and link from the MR. Small changes can stay in MR only.
Project Context & Rules: Preserving Collaboration Conventions
Maintaining context is knowledge work: deciding what info the current task needs, where to fetch it, and what old info should retire. Examples: test commands, special directory boundaries, tricky compatibility conditions — written into project instructions. Claude Code's official docs support .claude/rules/ with paths to scope applicability; these guide model behavior but don't replace hard constraints (tests, CI, permissions).
Practical difference:
"Watch compatibility when modifying interfaces" — hard to execute. "When modifying response fields of this interface, check registered legacy client contract tests; if breaking compatibility, submit migration plan first" — executable and auditable.
Rules should specify applicable objects, concrete actions, and supporting evidence. One-off exceptions shouldn't become universal rules.
Skills: Preserving Repeatable Methods
Turning useful flows into Skills accumulates procedural knowledge — "how to handle this class of task." Agent Skills official definition allows packaging instructions, scripts, reference materials, templates, loaded per task.
A reusable Skill typically defines trigger conditions, required inputs, execution steps, stop conditions, and success criteria. Copying a single successful conversation may retain too many accidental conditions.
Example: A release Skill with target environment check, change preview, execution steps, result verification, and failure handling has more reuse value than "help me release." If a flow succeeded only once, keep it experimental; revise during actual reuse.
Finding MCPs & Community Skills: Adopting External Capabilities
Valuable but distinct: "adopting a tool" vs. "forming project experience." MCP (Model Context Protocol) connects AI apps to external data, tools, workflows; it provides access channels. Knowledge lives in the connected Jira, doc stores, databases; specific servers may implement memory. The protocol itself doesn't auto-audit, organize, or maintain knowledge.
What's worth keeping: what problem the tool solves, why it fits the project, how to configure it, actual limitations, and what to watch when replacing it. Community Skills need checking for project paths, permissions, dependencies, success criteria. Tool exploration thus becomes reusable selection experience.
3. More Professional Approach: Let Different Records Carry Different Responsibilities
Engineering teams already have methods worth keeping; no need to reinvent for agents.
Architecture Decision Records (ADR) save "why." Michael Nygard's ADR format: context, decision, status, consequences (benefits & costs). Superseded decisions are kept and linked to replacements. Ideal for choices later developers might overturn without knowing original constraints.
Runbooks save "how." Fault diagnosis order, required permissions, stop conditions, recovery checks. Narrative explains reasons; scripts handle stable, repeatable operations; Skills describe how to select and invoke these materials.
Postmortems save "what went wrong and how to reduce recurrence." Google SRE emphasizes recording incident, impact, cause, preventive actions, and acknowledges postmortems have cost — need trigger conditions. Knowledge accumulation can be selective: repeated failures, major incidents, long-impact misjudgments deserve more curation.
Tests save "what behaviors must hold." Regression tests turn one bug into an executable check; contract tests record interface boundaries. Tests don't explain all context, so they should link bidirectionally with decision records.
These carriers have no fixed hierarchy. A short MR plus an effective test may suffice; major architectural changes need fuller decision rationale.
The author describes a workflow diagram (blue main path): task experience → review → reusable materials; feedback arrows return usage results to review. This is a recommended work loop, not an automated product mechanism.
Knowledge must be reviewed and actually reused to self-correct; obsolete material must exit.
4. Example: Turning One Task Record into Reusable Knowledge
Teaching example (not a real project, no measured data). Agent adds auto-retry to order interface. Review finds: request timeout may occur after order creation but before response returns; retry could create duplicate orders.
Steps to organize this finding:
Keep task evidence. Jira records phenomenon and acceptance criteria; MR links code changes, reproduction steps, test results. If risk is only inferred from code, mark as unverified — don't claim a production incident occurred.
Record applicability conditions. For this interface: when write result is uncertain, cannot judge failure by timeout alone. Auto-retry requires explicit idempotency support from the interface (e.g., agreed idempotency key, duplicate request detection & result return). Concrete implementation and lifecycle must be confirmed by the project.
Place guidance at the modification entry point. Order module rule: before modifying create-order retry logic, read linked decision and check duplicate-request regression cases. Don't over-generalize to "no HTTP request should ever retry."
Let tests constrain behavior. Tests cover agreed idempotent behavior (same request ID doesn't create second order) and project-defined edge cases (ID conflicts). Tests constrain business behavior; rules help agents find the check entry points.
Consider a Skill only after the process stabilizes. If multiple tasks repeatedly need write-retry review, extract a flow: identify side effects, check interface contract, find idempotency evidence, check failure scenarios, output review result. A one-off patch for a single interface may not warrant a dedicated Skill.
One experience yields multiple artifacts without full duplication. Task records link evidence; rules link decisions; Skills invoke processes and tests.
5. Avoiding Context Chaos When Saving Lots of Knowledge
Long-term storage can be rich; per-task loaded content must be curated. Anthropic's context engineering article discusses trimming to high-value context, keeping lightweight path references, and runtime on-demand retrieval. It warns that too many or vague tools increase decision paralysis. Practice: resident entry points provide few conventions and material locations; details loaded per task.
Diagram shows narrow entry: stored material can grow, but info entering current task must meet relevance and validity conditions. Right side retains "on-demand reading during execution" — don't interpret trimmed context as never reading details.
Small projects can start with an index: list material paths and one-line applicability notes per task ("modify order", "troubleshoot failed release"). Keyword search, directory structure, explicit links usually worth trying first. If cross-repo, cross-system retrieval becomes hard, then consider RAG (retrieval-augmented generation), which also needs handling permissions, provenance, updates, expiration.
Designate primary maintenance location per knowledge type: interface behavior → code + contract tests; current decisions → active ADRs; task history → Jira + MRs; rules → minimal guidance + links. This arrangement is a team agreement, not tool-enforced.
For long-lived materials, add maintainer, applicable version, last verification date, and define what changes trigger re-check. E.g., interface contract change → re-review related decisions; release command change → update runbook. Personal projects: maintainer can be yourself; no need for complex process.
Retain history and delete invalid content simultaneously: old decisions marked superseded, invalid ops removed from current entry points, duplicate instructions merged. Unsourced auto-summaries stay as pending-verification notes, not team rules.
6. How to Prove the Knowledge Is Useful
Most direct check: run a similar task in a new session or with an unfamiliar colleague. Observe: can they find relevant material, state applicability conditions, avoid the original error, meet acceptance criteria? If you must verbally prompt continuously, entry points or flows need improvement.
More systematic: keep a small set of representative tasks as an eval set. Anthropic's agent evaluation article discusses code checks, model review, human review; for coding agents, emphasizes clear tasks, stable environment, result testing.
Distinguish two things: regression tests verify business behavior correctness; agent evaluation also judges whether it found materials, chose the right process, and reduced human corrections. Code correctness ≠ effective knowledge entry points.
To compare before/after improvements, fix model version, task, environment, permissions; run multiple times. Record success rate, human correction count, time, cost. Model output varies; one success is only a signal. Without measurement, don't claim "efficiency improved by X%."
Personal use doesn't need an eval platform from day one. One cross-session reuse plus result verification already reveals many issues.
7. After Next Task, I'll Add These Steps
Prioritize curating changes that: fix recurring errors, reveal hard-to-see-from-code constraints, involve important trade-offs, or form stable processes. Ordinary small changes just need clear task descriptions.
For curation-worthy tasks, give the agent this closing prompt:
请检查本次任务是否产生了值得长期保存的知识。
优先考虑反复出现的问题、重要决策和稳定流程。
对于候选知识,请列出:
- 结论、适用条件,以及已知的不适用情况。
- 来源:代码、测试、日志或已确认的讨论。
- 未验证的假设,不要补写不存在的证据。
- 建议维护位置:任务记录、决策、操作手册、规则、Skill 或测试。
- 与现有材料的重复或冲突,以及需要更新的入口。
- 适用版本、维护者,以及哪些变化需要重新核验。
- 下一次怎样检查它是否仍然有效。
先提出可审核的修改。
仅在已授权范围内更新;待确认决定不得写成团队规定。
如果没有长期价值,说明理由即可。Then the author reviews facts, generalization scope, maintenance location. After approval, a follow-up task tests reuse effect. This prompt provides structure, not a substitute for review.
8. Back to the JD: How I'd Describe My Capability
If the hiring side wants to see reusable AI programming practices, current docs, rules, Skills are a valuable foundation. More compelling expression:
I use coding agents to organize task context and verification evidence, maintain long-lived conventions as project rules, distill repetitive flows into Skills, and verify reusability of these materials through linked tests and subsequent tasks.
This suits a capability statement. Resumes and interviews should only describe actually completed parts, ideally with a demonstrable example: original problem, artifacts left, how others found them, what error was avoided on reuse. If no effect data yet, state that validation methods are being established.
For the author, the next high-value step: pick a recently recurring problem, connect existing MRs, rules, or Skills, and run a new session to actually use them. This checks whether knowledge persisted, is accurate, findable, and worth the maintenance cost.
References
Verification date: 2026-10-11. Product docs update; confirm loading behavior with your version. The article's curation standards, teaching example, and implementation suggestions are the author's analysis, not official unified processes.
Claude Code: How Claude remembers your project
Agent Skills official overview
MCP: What is the Model Context Protocol?
Michael Nygard: Documenting Architecture Decisions
Google SRE: Postmortem Culture — Learning from Failure
Anthropic: Effective context engineering for AI agents
Anthropic: Demystifying evals for AI agents
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Ops Development & AI Practice
DevSecOps engineer sharing experiences and insights on AI, Web3, and Claude code development. Aims to help solve technical challenges, improve development efficiency, and grow through community interaction. Feel free to comment and discuss.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
