Mining users' AI chat sessions: gaps and forensics
- Overview
- From developer experience to agent experience
- Anatomy and distribution of a product skill
- What the research says about product skills
- Problems with product skills
- The docs-first approach
- Making docs accessible to agents
- Roles for tech writers with product skills
- Mining users' AI chat sessions: gaps and forensics
- Reimagining the documentation experience
Lesson 9 of 10
The previous topic, Roles for tech writers with product skills, placed identifying agent failures first on the roadmap. This topic covers how to gather that data. If your documentation portal includes an AI chat assistant, interaction logs provide direct evidence of where documentation falls short. The logs can reveal content gaps, vocabulary mismatches, and hallucinations in authentic user language.
As discussed in Evaluation suites aren’t user queries, test suites authored from internal specifications often fail to reflect authentic developer behavior. Chat logs provide the empirical basis for realistic evaluation sets. Without logs, decisions about product skills are often made in the dark, with eval tests that might provide artificially praising results about their efficacy. You need real data to evaluate whether skills are doing anything, and how to improve your docs.
Identifying documentation gaps in chat logs
Many documentation portals include conversational search assistants, and anonymized interaction logs are worth reviewing regularly. Transcripts show whether users received correct guidance, encountered errors, or abandoned sessions after failed queries. Without access to these logs, teams must guess where documentation leaves questions unanswered. Teams can also use automated scripts to process transcripts, flag unanswered queries, and identify topics needing coverage.
Analyzing user queries requires disciplined data governance. Anonymized chat sessions may still contain proprietary code snippets, private endpoints, or sensitive configuration details. Organizations must establish clear data policies, including automated redaction and defined retention windows.
Many modern documentation platforms provide analytics dashboards that surface these metrics natively, tracking unanswered questions and frequent failure points. Platforms such as Mintlify and Anthropic documentation provide interfaces for writers to review user queries and identify missing topics.
Reviewing logs also requires editorial judgment. An unanswered query might indicate an essential architectural gap, or it might represent an edge case that would add unnecessary clutter to the documentation. Determining which queries warrant new documentation remains a core editorial responsibility.
Sponsored
Natural user queries versus product feature lists
Session logs highlight the divergence between official product terminology and authentic developer language. Real user queries typically exhibit several recurring patterns:
- Queries describe high-level goals rather than naming specific features.
- Users reference legacy API terminology from older versions.
- Prompts omit critical environmental constraints, such as client-side execution boundaries.
- Users combine separate products into a single request without realizing they’re distinct tools.
Authoring evaluation prompts using internal product terminology tests only whether a model agrees with internal definitions. By contrast, chat logs provide evaluation data free from organizational naming/terminology bias.
This data informs two key documentation areas:
- Search terminology. Incorporating user phrasing into documentation headings, conceptual overviews, and skill frontmatter improves search matching.
- Architectural routing. Queries that span multiple products reveal where developers need explicit comparison guides, which forms the basis of The docs-first approach.
Documentation forensics and hallucination root causes
Documentation forensics involves investigating why an assistant provided incorrect guidance. When an assistant hallucinates, tracing the retrieval trajectory reveals where the breakdown occurred. Three primary causes account for most failures:
- Inaccurate source content. Outdated parameters or invalid code examples in existing documentation.
- Missing conceptual context. Unstated assumptions that force the model to guess procedures it can’t verify.
- Fragmented details. Information dispersed across disconnected pages that retrieval tools can’t assemble into a single context.
Forensic review requires direct access to chat logs and retrieval traces. On platforms that log only the final response without recording the retrieved passages, teams can see that an answer was incorrect but can’t diagnose why the retrieval failed.
Updating evaluation suites and documentation
Session logs directly inform ongoing documentation and skill maintenance. When a log reveals a content gap or an assistant hallucination, add that query to the evaluation suite using the user’s exact phrasing. Rewriting the query with formal product terminology removes the authentic phrasing that made the test case valuable. Once the query is recorded in the test suite, update the canonical documentation or skill, and run ablation tests to verify that the revision resolves the failure.
Chat logs identify where documentation fails real users, and evaluation suites confirm whether revisions resolve those failures. Together, they turn documentation maintenance into an empirical, test-driven process.
Continue to the next topic: Reimagining the documentation experience
About Tom Johnson
I'm an API technical writer based in the Seattle area. On this blog, I write about topics related to technical writing and communication — such as software documentation, API documentation, AI, information architecture, content strategy, writing processes, plain language, tech comm careers, and more. Check out my API documentation course if you're looking for more info about documenting APIs. Or see my posts on AI and AI course section for more on the latest in AI and tech comm.
If you're a technical writer and want to keep on top of the latest trends in the tech comm, be sure to subscribe to email updates below. You can also learn more about me or contact me. Finally, note that the opinions I express on my blog are my own points of view, not that of my employer.