---
title: "Podcast: Doc testing, skills files, and the guardians of knowledge – with Manny Silva"
date: 2026-03-08
description: "Our conversation in this podcast covers some of the topics he’s exploring. For example: documentation testing (testing docs vs. testing the product), skills files (versus regular markdown files that don’t..."
canonical_url: https://idratherbewriting.com/blog/podcast-silva-guardians-of-knowledge
---
# Podcast: Doc testing, skills files, and the guardians of knowledge – with Manny Silva
> In this podcast, Fabrizio Ferri-Benedetti ([passo.uno](https://passo.uno)) and I chat with Manny Silva ([instructionmanuel.com](https://instructionmanuel.com)), head of documentation at Skyflow and author of [Docs as Tests](https://www.amazon.com/Docs-Tests-Resilient-Technical-Documentation/dp/0994169361). Manny is working on a follow-up book that incorporates AI, covering validated generation, trusted agents, and self-healing documentation.

 Our conversation in this podcast covers some of the topics he’s exploring. For example: documentation testing (testing docs vs. testing the product), skills files (versus regular markdown files that don’t follow the skills spec), the consultant model of docs (and whether this is the future of tech comm in companies), externalizing and sharing skills files (and why one might or might not want to do that), and much more. Throughout, we wrestle with the big question lurking behind all of it: as tech writers pour their expertise into systems that machines can run, are we accelerating ourselves or automating ourselves out of a job?

 - [Transcript](#transcript)

## Video

## Audio only

**Listen here:**

[![](https://s3.us-west-1.wasabisys.com/idbwmedia.com/images/apple_podcasts.png)](https://itunes.apple.com/us/podcast/id-rather-be-writing-podcast/id277365275)

[![](https://s3.us-west-1.wasabisys.com/idbwmedia.com/images/watchonyoutubeblack.png)](https://www.youtube.com/@idratherbewriting)

[![](https://s3.us-west-1.wasabisys.com/idbwmedia.com/images/spotify.png)](https://open.spotify.com/show/4HeOZfPGMMfViOhVS40QBD)

## Resources mentioned

 - 
[The Most Actionable Docs Around: Agent Configs](https://instructionmanuel.com/agent-configs-are-docs) (Manny Silva)

 - 
[Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?](https://arxiv.org/abs/2602.11988) (Gloaguen et al.)

 - 
[Docs as Tests: A Strategy for Resilient Technical Documentation](https://www.amazon.com/Docs-Tests-Resilient-Technical-Documentation/dp/0994169361) (Manny Silva)

 - 
*Docs as Tests and AI: Validated Generation, Trusted Agents, and Self-Healing Technical Documentation* (forthcoming book Manny is currently working on – not yet released)

 - 
[Skills are docs, and docs need tech writers](https://passo.uno/skills-are-docs/) (Fabrizio Ferri-Benedetti)

 - 
[Write the Docs Portland 2026](https://www.writethedocs.org/conf/portland/2026/), [Writing Day](https://www.writethedocs.org/conf/portland/2026/writing-day/)

## Topics covered in this podcast

Here’s a list of topics we talked about. (Note: AI-generated.)

 - 
 **Doc testing vs. product testing** — Manny draws a clear line between testing your documentation’s accuracy (can a user follow these steps?) and testing the product itself (does the API actually work?). The key takeaway: figure out the minimum viable thing you can test in your docs and let QA be QA.

 - 
 **Manny’s doc setup and workflow** — Manny maintains two doc sets (Doc Detective and Skyflow), runs the full test suite on every PR, and has scheduled daily runs at midnight. Tests cover UI procedures, CLI commands, API calls, automated screenshots via visual regression, and even video recording.

 - 
 **Deterministic vs. probabilistic testing** — Deterministic tests (like Doc Detective in CI/CD) give the same answer every time, which is the signal Manny’s looking for. Probabilistic tools like browser use can bootstrap those tests but can’t be trusted on their own — ensemble testing helps but gets expensive fast.

 - 
 **Skills files vs. personal instruction files** — Tom has ~40 markdown instruction files he drags into Gemini but hasn’t formalized them as skills. Manny and Fabrizio argue that the skills spec makes them portable across agent harnesses (Claude Code, Gemini CLI, Codex), shareable with teams, and accessible for autonomous agent workflows.

 - 
 **Writing skills manually vs. AI-assisted reflection** — Tom describes a loop where he has AI attempt a task, reviews its thought log for friction points, then has the AI update the instruction file. Manny validates this approach — the human doesn’t have to do the literal writing, but must curate what goes in or gets cut.

 - 
 **The ethical dilemma of externalizing knowledge** — Tom poses the uncomfortable question: if I pour all my expertise into skills files that anyone (or any machine) can run, am I automating myself out of a job? Manny’s answer: own it or someone else will. Security by obscurity won’t last when an LLM can find your half-documented process and scaffold it into a skill.

 - 
 **The consultant model vs. the curator model** — Manny distinguishes two futures for tech writers. The consultancy model: you parachute in, set up docs, and leave with no ongoing quality bar. The central pillar model (what he’s built at Skyflow): you own all documentation, accept contributions from anyone or anything, but maintain the authority to block publication.

 - 
 **Guardians of knowledge and content curation** — Fabrizio envisions tech writers as librarians guarding high-quality, human-made knowledge that machines shouldn’t touch — the “high signal” that feeds MCP servers, skills, and downstream AI tools. Manny prefers the term “content curator”: not gatekeeping, but holding standards and verifying accuracy through testing.

 - 
 **Fabrizio on skills as documentation** — Fabrizio connects skills, agents.md, and agent definitions back to what tech writers already do — it’s all documentation. He notes that companies with the best-curated knowledge repositories will have the best AI outcomes, and that tech writers are surprisingly good at prompting LLMs compared to developers and PMs.

 - 
 **Manny’s book and Write the Docs Portland** — Manny’s writing a follow-up to *Docs as Tests* that incorporates AI, evals, agentic workflows, and self-healing documentation systems. He’ll also be hosting a writing day table at Write the Docs Portland 2025, where attendees can bring their docs to get tested with Doc Detective.

![Podcast: Doc tests, Skills, and the guardians of knowledge -- with Manny Silva](https://s3.us-west-1.wasabisys.com/idbwmedia.com/images/episode5mannysilvapodcast.png)

## Shorts

Here are some shorts pulled from the longer video.

## Narrative summary

If the podcast were in the form of an article, this is what it would look like. (Note: AI-generated.)

### The Tech Writer’s New Job: Building the Machine That Builds the Docs

There’s an uncomfortable question hanging over technical writing right now, and it goes something like this: if I pour everything I know into a system that a machine can run, what exactly am I still needed for?

It’s the kind of question that tends to produce either breezy optimism (“AI is just a tool!”) or existential dread. But a recent conversation between Tom Johnson, Fabrizio Ferri-Benedetti, and Manny Silva — three tech writers working at the bleeding edge of documentation and AI — suggests that the real answer is more interesting than either of those poles. The tech writer’s job isn’t disappearing. It’s mutating into something that looks less like writing and more like orchestration.

### From Testing Docs to Testing Everything

Manny Silva, head of documentation at Skyflow, has built something most tech writers only talk about: a fully automated documentation testing pipeline. On every pull request, a suite of deterministic tests verifies that the procedures in his docs actually work. A tool called Doc Detective opens a browser, clicks through UI steps, runs CLI commands, and compares screenshots against previous versions. If a button moved, if an API response changed, if a procedure no longer works — the CI check fails.

This isn’t QA. Manny is careful to draw that line. He’s not testing whether the product works; he’s testing whether the documentation’s claims are accurate. The distinction matters, because it reframes what tech writers are responsible for. Not the product. The knowledge layer on top of it.

But here’s where it gets interesting. Testing docs for accuracy is only the first step. Manny’s upcoming book extends the same testing philosophy to a new category of documentation that barely existed two years ago: skills files, agent definitions, and agentic workflow instructions. These are the markdown files that tell AI tools how to perform tasks — how to write release notes, how to debug a test failure, how to follow your team’s conventions. And just like traditional docs, they can be wrong. They can drift. They can silently degrade the quality of everything downstream.

So you test them too. You write evals — evaluations that check whether an AI agent’s output meets your criteria. You define entry and exit conditions for each task. You build feedback loops. The tech writer becomes the person who ensures the entire knowledge system is trustworthy, from the docs a human reads to the instructions a machine follows.

### The Skill File Dilemma

Tom Johnson raised what might be the sharpest version of the existential question. He has about 40 markdown instruction files that walk AI tools through his release notes process — generating reference docs, summarizing file diffs, tying changes to roadmap items, updating API diagrams. Each one has been refined through an iterative loop: run the task, review the AI’s thought log for friction points, update the instructions, repeat.

These files represent years of accumulated expertise compressed into executable form. And Tom’s question was blunt: why would I formalize these into shareable skills files that anyone — or any machine — can run without me?

Manny’s answer was equally blunt: because if you don’t, someone else will. An LLM trawling your codebase will eventually find your half-documented process, scaffold it into a skill, and hand it to a PM who doesn’t know the difference between your carefully curated version and the auto-generated one. The choice isn’t between sharing your knowledge and keeping it. It’s between owning the system and having the system built without you.

This is the core argument: tech writers who externalize their expertise into testable, maintainable skill files aren’t automating themselves out of a job. They’re claiming ownership of a new and increasingly critical layer of infrastructure. The alternative — hoarding knowledge in personal files and hoping nobody notices — is security by obscurity, and it won’t survive the first curious AI agent.

### Curators, Not Gatekeepers

The three writers converged on a term for what this role actually looks like: content curator. Not gatekeeper — Manny emphasized that he actively wants contributions from engineers, PMs, even AI agents. But someone has to hold the quality bar. Someone has to decide what gets published and what gets blocked. Someone has to verify that the knowledge feeding your MCP servers, your chatbots, your agentic workflows is accurate.

Fabrizio put it more romantically: guardians of knowledge, like librarians guarding the old tomes. But the practical version is less poetic and more essential. Companies that maintain strongly curated, well-tested repositories of high-quality knowledge will have better AI outputs than companies that don’t. The signal-to-noise ratio of your documentation directly determines the signal-to-noise ratio of everything built on top of it.

Engineers don’t want this responsibility. They never have. They didn’t want it when wikis were supposed to democratize documentation, and they don’t want it now that AI can draft a passable doc in seconds. What they want is someone to make sure it’s right. That’s always been the job. The tools have changed. The accountability hasn’t.

The tech writer’s future isn’t writing. It’s building, testing, and curating the systems that write — and making sure the knowledge those systems run on is worth trusting.