Search results

For AI agents: a documentation index is available at https://idratherbewriting.com/llms.txt. Markdown versions of all pages are available by appending .md to any page URL.

Making docs accessible to agents

Last updated: Sep 30, 2026 comments

The docs-first approach rests on an assumption that’s easy to skip past. If agents are supposed to get their guidance from your documentation, they have to be able to read it. However, many documentation sites are built in ways that work well in a browser and poorly for an agent. What does an agent actually see when it fetches one of your pages? Does it see the content behind your tabs? Does it get past your bot protection, or does it get a challenge page instead? And if your pages load their content with JavaScript, does the agent see anything at all?

This topic walks through how agents read documentation pages, the ways pages become hard or impossible for agents to read, and what to do about each one. From developer experience to agent experience introduced Markdown mirrors and /llms.txt as delivery layers. This topic covers the practical work of making those layers function, along with the other problems that keep agents from reading a page.

How agents read a page

Most coding agents don’t read your page in a browser. They make an HTTP request, convert the HTML response to Markdown or plain text, and cut the result off at a size limit. Many then pass what’s left to a smaller model that answers whatever the agent asked about the page, so the agent works from a summary shaped by its own question rather than from the page itself (Shilkov). These fetch tools generally don’t run JavaScript or click on anything. Anthropic’s web fetch tool, for example, doesn’t support pages rendered with JavaScript (Anthropic).

The details vary by tool, which is part of what makes this work tricky:

  • Claude Code converts HTML to Markdown with the Turndown library, truncates the result at about 100 KB of text, and passes it to a smaller model before Claude sees it (Shilkov). Only certain trusted sites that serve Markdown under that limit skip the extra step (Gurgone).
  • Cursor asks for Markdown through content negotiation (Checkly). Its size limits range from about 28 KB to more than 240 KB, depending on which fetch method it picks (Rodriguez).
  • GitHub Copilot’s fetch tool returns relevance-ranked excerpts, averaging about 13,000 characters, rather than whole pages (Rodriguez).
  • The reference MCP fetch server cuts content off at 5,000 characters by default, although the model can ask for a higher limit or for later chunks (MCP fetch server).

In other words, the same page might arrive nearly whole in one tool, as excerpts in another, and as its first 5,000 characters in a third. Agents also tend not to notice when content is missing. In testing, agents sometimes treated filtered excerpts as the complete page, because the excerpts read coherently on their own (Rodriguez).

Several fixes in this topic overlap with web accessibility practices, such as text alternatives for images and descriptive headings. The overlap isn’t complete, though. Screen readers work from the page a browser has already rendered, so they see JavaScript-generated content that most agents miss. The same fixes also help any documentation chatbot or MCP server that indexes your published site, since its ingestion step fetches pages much the way an agent does.

Content that isn’t in the HTML

Some failures keep content from reaching the agent at all. These are the most damaging problems in this topic, and they’re easy to miss, because the page looks fine to anyone viewing it in a browser.

Client-side rendering. If your docs are a single-page app that builds each page in the browser, the HTTP response contains framework code, styles, and navigation, but no documentation. The server still returns a 200 status, so the agent has no reason to suspect a problem. It tries to work from the navigation links or falls back on its training data.

In one set of controlled probes, a client that didn’t run JavaScript couldn’t retrieve an answer that existed only in a JavaScript payload (Vercel). The framework itself usually isn’t the cause, since most modern frameworks can render pages on the server. To fix the problem, turn on server-side rendering or pre-rendering for documentation pages. If only some templates load content in the browser, such as a page whose code samples change with a language picker, you can fix those templates instead of rebuilding the whole site. API reference explorers that render endpoint documentation in the browser from an OpenAPI file have the same problem, often on the pages agents need most, so pre-render those as static HTML too.

Content that loads on interaction. Tabs or accordions that fetch their content when someone clicks, “show more” buttons, infinite scroll, and pages reachable only through a search box all hide content from agents, because agents don’t click or type. Content that’s in the HTML but hidden with CSS is a different case. Converters ignore CSS, so agents usually do get that content, which causes the tab problem described in the next section.

Images, embeds, and tooltips. Converters keep an image’s alt text and URL, not the image itself, so a diagram or a screenshot of code reaches the agent as whatever the alt text says. Code samples embedded through JavaScript or iframes often don’t arrive at all, and neither do definitions that appear only in hover tooltips. Write code as text rather than screenshots, give diagrams alt text that states what they show, link videos to transcripts, and put tooltip definitions in the page text or a glossary.

Content that arrives in an unusable form

Other failures deliver the content, but in a form that pushes the important parts out of reach or scrambles their structure. The fixes here are mostly about page design rather than infrastructure.

Boilerplate before the content. Navigation menus, sidebars, and inline CSS or JavaScript count against an agent’s size limit just like your prose does. Turndown, the converter Claude Code uses, doesn’t remove any elements by default, and it outputs the text of any tag it has no rule for, so the contents of an inline <style> block can come through as raw CSS (Turndown). The overhead adds up quickly. Cloudflare measured one of its own blog posts at 16,180 tokens as HTML and 3,150 tokens as Markdown, an 80% reduction (Cloudflare). Move CSS and scripts into external files, which agents generally don’t fetch, and keep the markup before your main content short.

Long pages. Anything past an agent’s limit is gone, and the agent often doesn’t notice. Limits vary widely, from 5,000 characters for the MCP fetch default to about 100 KB for Claude Code, so a page that fits comfortably in one tool can get cut off in another. Split long tutorials and reference pages into smaller pages that each make sense on their own. Avoid slicing one topic into numbered windows, though, since an agent that gets one window doesn’t have a complete answer.

Tabs and dropdown filters. Because hidden tab panels are still in the HTML, a page with eight language tabs converts into one long stream that contains every variant in source order. An agent might see only the first few variants before the cutoff, and asking for a specific language doesn’t help if that tab falls past it. (In one test of a MongoDB tutorial by Dachary, a request for the Python version came back with no Python at all.) Repeated headings inside tabs, such as “Step 1” and “Step 2,” make the problem worse, since nothing tells the agent which variant a step belongs to. Put each major variant on its own page, or at least add the variant to each heading (“Step 1 (Python)”), and put the most commonly used variant first.

Tables. Tables are mostly fine, with some caveats. When you serve Markdown yourself, simple tables come through intact, because Markdown has a table syntax. When the agent converts your HTML instead, the result depends on its converter. Turndown’s default rules don’t include tables (table support comes from a separate plugin), so a converter running the defaults turns each cell into its own paragraph and loses the rows and columns (Turndown). Even a good converter can’t express merged cells, or lists and code blocks inside cells, because Markdown tables can’t contain block-level elements (GFM spec).

Very large generated tables cause a different problem, since they can push the rest of the page past an agent’s limit. To keep tables usable, serve Markdown, keep tables to simple rows and columns, state critical facts such as constraints in prose rather than only in a table cell, and put explanatory prose before large tables so that truncation removes rows rather than explanation.

Sponsored

Pages the agent can’t reach

The failures in this section stop an agent before it reads anything. They also tend to be invisible from the inside, since the people who run a docs site rarely browse it the way an agent does.

Bot protection. CDN bot management and firewall rules tuned for scrapers often can’t tell an agent fetching docs for a developer from abuse. In the same controlled probes mentioned earlier, a server that returned 403 to agent user agents blocked both a plain fetch client and one that ran JavaScript (Vercel). These failures are often hard to spot. A challenge page or rate limit that a person clicks through or never notices can stop an agent cold, and some limits start only after the first few requests, so a quick check in a browser looks fine while a multi-page agent session fails. Exempt documentation routes from aggressive bot enforcement. Where you do need limits, return a 429 status with a Retry-After header rather than a silent challenge page, which gives the agent a clear signal to back off (Vercel). Test by fetching several pages in a row rather than one.

robots.txt rules. Blocking AI crawlers such as ClaudeBot or GPTBot in robots.txt affects training crawlers, which identify themselves. It mostly doesn’t affect coding agents, because many of them send generic or browser-like user-agent strings. In one test, Claude Code identified itself as axios/1.8.4 and Cursor as a version of Chrome (Checkly). A few agents do identify themselves, such as OpenAI’s Codex, which sent ChatGPT-User in the same test, so a broad rule against AI user agents can still catch them. Decide about training crawlers and agent fetches separately.

Login walls. An agent that hits a login wall gets a 401 or 403 error, a login page served with a 200 status, or a redirect to a single sign-on provider on another host. In each case, it falls back on training data or goes looking for blog posts about your product, sometimes without telling the user. If you must gate some documentation, keep the reference docs public, publish a public /llms.txt that describes what exists, ship docs with your SDK, or offer an MCP server that handles authentication on the agent’s behalf.

Moved and broken URLs. Agents often guess at URLs, and without a map of the site, they request pages that don’t exist (Mintlify). Moved content makes this worse, because a URL that an agent remembers from training might no longer work. Same-host redirects with a 301 status work well, since the HTTP client follows them without the agent noticing. Claude Code doesn’t automatically follow redirects to a different host. It reports the new URL instead, and the agent has to make a second request (Shilkov).

JavaScript redirects don’t work at all for a client that doesn’t run JavaScript. Soft 404s, meaning error pages served with a 200 status, are worse than real 404s, because they remove the signal that tells the agent a page doesn’t exist (Vercel). Keep URLs stable, redirect on the same host when you must move content, and return a real 404 for missing pages.

Serve Markdown versions of your pages

Serving Markdown sidesteps most of the conversion problems above, because the agent gets clean content instead of whatever its converter makes of your HTML. Fix the HTML first, though, since many agents never ask for Markdown. Serve the Markdown in two ways:

  • .md URLs. Make each page available at its normal URL with .md appended, such as /guide/authentication.md. Agents use these URLs when llms.txt or a note on the page tells them the URLs exist.
  • Content negotiation. When a request includes Accept: text/markdown, return the Markdown version from the normal page URL, with a Content-Type: text/markdown header and a Vary: Accept header so that caches keep the two versions separate (MDN). In one test, Claude Code, Cursor, and OpenCode requested Markdown this way, while Codex, Gemini CLI, Copilot, and Windsurf didn’t (Checkly). In a larger study, 65% of agent fetches asked for Markdown (Vercel). The Markdown is usually much smaller than the HTML, by 74% in one measurement, and the cache setup matters, because a misconfigured cache can serve Markdown to a person’s browser (Pandiyan).

How you build the Markdown depends on your tooling:

  • Documentation platforms. Many hosted platforms generate Markdown versions automatically. Mintlify generates .md versions, /llms.txt, and /llms-full.txt for the sites it hosts (Mintlify). Read the Docs answers Accept: text/markdown requests on its hosted sites with no configuration (Read the Docs). Fern generates both index files as part of the docs build (Fern). Check whether your platform does this, and whether the feature is turned on.
  • Static site generators. Generate a Markdown version of each page at build time from the same source as the HTML, so the two versions can’t drift apart. The llms.txt site lists plugins for several generators and content management systems.
  • Your CDN. Cloudflare’s Markdown for Agents converts HTML to Markdown at the edge when a request asks for it, with no code to write. It’s available on Pro, Business, and Enterprise plans, strips navigation, headers, footers, scripts, and styles, and handles pages up to 2 MB. Unless your server sets its own Content-Signal header, Cloudflare adds one that tells crawlers your content can be used for AI training, which you might or might not want.
  • A worker or edge function. On this site, I added a Cloudflare worker in about five minutes that serves the Markdown version of any page, both when you append .md to the URL and when a request asks for Markdown.

Check the Markdown you serve

A Markdown generator is a second rendering pipeline, and it needs its own quality checks. People look at your HTML every day, but almost nobody reads the Markdown, so it can stay broken without anyone noticing. That happened on this site. My worker stopped converting each page at the first ad block in the article, so agents that asked for Markdown got only the content before that ad, often less than a fifth of the page. The skills overview, for example, came through as 382 of its roughly 1,980 words. Nothing looked wrong in a browser, and the problem only showed up when the word counts of the Markdown and HTML versions were compared. Check for these problems:

  • Missing content. A Markdown pipeline can drop sections, stop partway through a page, or skip pages entirely while the HTML looks fine. Compare the Markdown and HTML versions of a sample of pages, and check the last section of each.
  • Unclosed code fences. An unclosed code fence turns everything after it into code, so the agent reads the rest of the page as literal text rather than as instructions.
  • Relative links. Agents often lose track of a page’s original URL once its content passes through a summarizing model or gets split into chunks, and a link like /guide/auth means nothing without it. Use absolute URLs in the Markdown you serve.
  • Wrong content types. Make sure each .md URL returns Markdown with a text/markdown content type, not an HTML error page with a 200 status.

Help agents find the better versions

Agents rarely discover a Markdown mirror or an /llms.txt file on their own. They mostly find these files through links. In one study, 86% of agent fetches of llms.txt came through a link rather than a guessed path, and about a third of the runs that reached the file went on to use a page it listed (Vercel). Mintlify’s benchmark showed what happens once agents do find the file, with a link to /llms.txt cutting their 404 errors to near zero. Content negotiation is the exception to the discovery problem, since agents that send Accept: text/markdown get Markdown without knowing anything about your site. To help agents find the rest:

  • Publish an /llms.txt index. The format is a Markdown file with an H1 heading, a short summary in a blockquote, and lists of links (llms.txt proposal). Link to the Markdown version of each page with absolute URLs, and keep the file small enough to fit in a single fetch. For a larger site, use a short top-level index that links to section-level index files. Point the index at your current release, so agents skip outdated versions (GitBook). Some platforms also generate an /llms-full.txt that concatenates every page into one file, which suits APIs small enough to fit in a context window (Fern).
  • Point to it from every page. Add a one-line note at the top of each page, in both the HTML and the Markdown, that tells agents where to find the index and the Markdown versions. The documentation from Anthropic, Cloudflare, and Payabli all includes a note like this. You can also advertise each page’s Markdown version with a <link rel="alternate" type="text/markdown"> tag in the page’s <head> (Vercel). Some sites hide the note visually and leave it in the HTML for converters. If you do, keep in mind that screen readers still announce visually hidden text.

How to test your docs

You can check most of these problems in a few minutes without special tools. Pick a handful of representative pages, such as a long tutorial, a page with tabs, an API reference page, and a page with a large table, and then do the following:

  1. Fetch a page without a browser, and search for a phrase that appears near the end of the page in your browser. (Pick a phrase without quotation marks or apostrophes, since the HTML might encode them differently.)

    curl -s https://docs.example.com/guide/auth | grep -c "a phrase near the end of the page"
    

    A count of 0 means the phrase isn’t in the HTML the agent receives.

  2. Request the Markdown version both ways, and check that each response starts with your content rather than navigation:

    curl -s -H "Accept: text/markdown" https://docs.example.com/guide/auth | head -40
    curl -s https://docs.example.com/guide/auth.md | head -40
    
  3. Check the response headers for Content-Type: text/markdown and Vary: Accept:

    curl -sI -H "Accept: text/markdown" https://docs.example.com/guide/auth
    
  4. Compare the Markdown with the HTML. Check the last section of the page, the last tab, and the last rows of any large table.

  5. Ask your agent to fetch the page and quote its final paragraph, or a sentence from the last tab. Then check the quote against the page. Agents sometimes report that they read a whole page when they didn’t.

  6. For a fuller report, run npx afdocs check https://docs.example.com, which checks a site against the Agent-Friendly Documentation Spec. To see how your own agent handles these failures, point it at agentreadingtest.com.

Where to start

If you can only do a few things, start with the problems that block agents entirely. Make sure your content is in the HTML, and check that bot protection lets agents through, since nothing else matters when an agent gets an empty shell or a challenge page. After that, serve Markdown and check it against the HTML, publish an /llms.txt index and point to it from every page, keep pages small, and keep URLs stable.

Most of this work is configuration rather than writing, which makes it some of the cheapest work in this chapter. A reasonable first step is to run the six checks above against your five most-visited pages and note what’s missing.


Continue to the next topic: Roles for tech writers with product skills

Comment on LinkedIn

About Tom Johnson

Tom Johnson

I'm an API technical writer based in the Seattle area. On this blog, I write about topics related to technical writing and communication — such as software documentation, API documentation, AI, information architecture, content strategy, writing processes, plain language, tech comm careers, and more. Check out my API documentation course if you're looking for more info about documenting APIs. Or see my posts on AI and AI course section for more on the latest in AI and tech comm.

If you're a technical writer and want to keep on top of the latest trends in the tech comm, be sure to subscribe to email updates below. You can also learn more about me or contact me. Finally, note that the opinions I express on my blog are my own points of view, not that of my employer.