Building Agent-Friendly Documentation Sites: Rspress's llms.txt, SSG-MD & Markdown Negotiation
This article explains how Rspress makes documentation sites accessible to AI agents through llms.txt indexes, SSG-MD for generating Markdown alongside HTML, Accept: text/markdown content negotiation, AFDocs compliance checking, and injectLlmsHint for discoverability, achieving a perfect 100/100 AFDocs score.
llms.txt: The Sitemap for the Agent Era
The llms.txt specification defines a Markdown index placed at a site's root or subpath, containing a site description and links to detailed content. The only required element is a level-1 heading for the project name; optional sections include a blockquote summary, supplementary notes, grouped link lists under level-2 headings, and an Optional group for content that can be skipped when context is limited. Rspress generates three artifact types in doc_build/: HTML pages for humans, Markdown pages for agents, and index files ( llms.txt for a concise index, llms-full.txt for the full site Markdown). While sitemap.xml targets search-engine crawlers, llms.txt enables progressive disclosure for agents — first a lightweight index, then on-demand page reads.
SSG-MD: Static Site Generation to Markdown
Rspress introduces Static Site Generation to Markdown (SSG-MD), a first-class capability alongside traditional SSG. The analogy table contrasts the two:
Full name: Static Site Generation vs. Static Site Generation to Markdown
Optimization target: SEO (search-engine optimization) vs. GEO (generative-engine optimization)
Target audience: Search-engine crawlers vs. LLMs / vector-retrieval systems
Index file: sitemap.xml vs. llms.txt Full-content file: (none) vs. llms-full.txt Core implementation: renderToString vs. renderToMarkdownString Access pattern: /guide/start/introduction.html vs.
/guide/start/introduction.mdWhy SSG-MD Is Needed
In React-based frameworks (including MDX), dynamic components make static information hard to extract. Feeding raw MDX to AI introduces code syntax noise and loses React component output; converting HTML to Markdown often yields poor fidelity. SSG solves this for SEO by emitting static HTML; SSG-MD solves it for GEO by emitting high-fidelity Markdown directly from the virtual DOM during render.
How SSG-MD Is Implemented
renderToMarkdownString: Rspress implements a renderToMarkdownString method (similar to react-dom 's renderToString) that renders React components to Markdown strings. This API is generic and published as react-render-to-markdown.
remarkSplitMdx plugin: A custom Remark plugin splits the MDX AST, serializing pure Markdown text as string literals while preserving JSX components and MDX expressions (e.g., {variable}) as React elements. This ensures Markdown passes through untouched, and dynamic components are rendered by renderToMarkdownString.
Environment flag: import.meta.env.SSG_MD lets components differentiate SSG-MD from browser rendering. Example: a Tab component returns Markdown-formatted text ( **Here is a Tab named ${label}**) during SSG-MD, but a normal <div> in the browser.
Internal component adaptation: Rspress built-ins like <PackageManagerTabs command="create rspress@latest" /> render appropriate Markdown during SSG-MD (shown in the article's screenshot).
Accept: text/markdown — Content Negotiation for Markdown
In November 2025, Claude Code lead Boris Cherny announced that Claude's WebFetch tool automatically sends Accept: "text/markdown, *". Cloudflare observed the same header from Claude Code, OpenCode, and other coding agents. Compared to HTML, Markdown omits navigation, styles, and scripts, reducing token usage and eliminating the need for agents to extract main content from HTML. Rspress, being a pure static framework, has no deploy server to negotiate responses; instead, the Rspress website configures a Cloudflare rewrite rule: when the request contains Accept: text/markdown, the path is internally rewritten to the corresponding .md file, while browsers still receive HTML. A curl example demonstrates the Markdown response, which includes a hint pointing to llms.txt.
AFDocs: Verification and Scoring
AFDocs (Agent-Friendly Documentation Spec) is an open-source checker with 7 categories and 23 checks. The article maps Rspress features to AFDocs checks: llms.txt → llms-txt-exists, llms-txt-valid (can agents find and parse the index?)
SSG-MD → markdown-url-support (does every page provide Markdown?) Accept: text/markdown → content-negotiation (does the header return Markdown?) injectLlmsHint → llms-txt-directive-html, llms-txt-directive-md (can agents discover the index from a single page?)
Dual artifacts → markdown-content-parity (do HTML and Markdown express the same content?)
Rspress scores 100/100 on AFDocs. The command npx afdocs check https://docs.example.com --format scorecard produces a detailed report with results, suggestions, and scores for before/after comparison. Mintlify's Agent Score builds on the same spec, adding checks for full content, Agent Skills, and MCP Server discoverability. AFDocs does not assess article quality or answer correctness — only verifiable rules.
injectLlmsHint: Exposing llms.txt Location
AFDocs' Content Discoverability checks require that agents landing on any HTML page know about llms.txt and the current page's Markdown, and that agents receiving a Markdown page know where to find the full site index. Rspress's injectLlmsHint config injects a visually hidden
<div class="rp-llms-hint" style="position:absolute;width:1px;height:1px;padding:0;margin:-1px;overflow:hidden;clip:rect(0,0,0,0);clip-path:inset(50%);white-space:nowrap;border:0">containing plain-text URLs for llms.txt, llms-full.txt, and the current page's Markdown. This avoids display:none, hidden, or aria-hidden so the hint remains in the DOM and survives HTML-to-Markdown conversion. The Markdown artifact uses a blockquote format ( > For AI agents: ...) listing only llms.txt and llms-full.txt (since the current page is already Markdown). AFDocs explicitly allows clip-rect or sr-only visual hiding but requires the directive to stay in the DOM and survive conversion.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
ByteDance Web Infra
ByteDance Web Infra team, focused on delivering excellent technical solutions, building an open tech ecosystem, and advancing front-end technology within the company and the industry | The best way to predict the future is to create it
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
