how to document tribal knowledge for agents
Documents tribal knowledge so agents and engineers can find it: decision logs for the why, runbooks for the how, incident writeups for the what-happened, in one versioned repo with consistent naming, review, and expiry. Use when answers live only in chat history or senior engineers' heads. Not for fast-changing processes or schema documentation.
TL;DR
Agents are only as good as the knowledge they can find. Capture tribal knowledge as short, searchable, versioned docs: decision logs for the why, runbooks for the how, incident writeups for the what-happened. One repo, predictable paths, reviewed like code. If the answer only exists in someone's head, the agent cannot use it and neither can the next hire.
Error / query
how to document tribal knowledge for agentsUse this skill when
- the same questions get asked in chat every month
- an agent gives wrong answers because the real knowledge is unwritten
- a senior engineer is leaving and their knowledge is not written down
- you are building the knowledge base an SRE agent will search
Not for this skill when
- the knowledge changes daily (stabilize the process first)
- you need a data catalog or schema docs (different effort)
- the audience is only humans who already know the context
Steps
Step 1: Create the knowledge repo with sections agents can navigate
mkdir -p knowledge/decisions knowledge/runbooks knowledge/incidents knowledge/glossary
ls knowledge/Expected: the four directories. Decisions hold the why, runbooks the how, incidents the what-happened, glossary the terms. Agents navigate structure better than search alone, so the layout is part of the documentation.
Step 2: Write decision logs as short dated records
ls knowledge/decisions/ | tail -5
cat knowledge/decisions/2026-09-15-why-we-pinned-kafka-version.mdExpected: one file per decision with date, context, options considered, and the chosen option with reasons. Five paragraphs max. The decision log answers the question agents get wrong most: "why is it like this", which no amount of code reading reveals.
Step 3: Make every doc findable with consistent naming and an index
grep -r -l -i "failover" knowledge/ | head -10
cat knowledge/index.md | head -30Expected: grep finds the relevant docs and the index lists them by topic. File names carry dates and topics; the index is hand-curated, not generated, because curation is what makes it trustworthy.
Step 4: Review knowledge changes like code changes
gh pr create --title "knowledge: document payments retry policy" --body "source: incident INC-2026-0912"
gh pr checks [PR_NUMBER]Expected: the PR reviewed by someone who knows the topic. Unreviewed knowledge rots into folklore; the review bar is "would I trust an agent to act on this at 3am".
Step 5: Expire and refresh on a schedule
find knowledge/ -name "*.md" -mtime +180 | head -20Expected: the list of docs untouched in six months. Each gets one of three fates: confirmed still true, updated, or deleted. Stale knowledge an agent treats as current is worse than no knowledge.
Variant phrasings
"knowledge base for ai ops agent"
Same repo, plus machine-readable front matter on each doc (service, topic, last-verified date) so the agent can filter by freshness and scope.
"how to capture knowledge before someone quits"
Interview them against the incident list: for each major incident they handled, write the decision log and runbook that did not exist. Two weeks of this beats a year of "write down what you know".
"wiki vs git for team knowledge"
Git wins when agents consume it: version history, review, diff, and blame. Wikis win for casual collaboration; use git for anything an agent will act on.
Why it happens
Tribal knowledge stays tribal because writing it down has no immediate payoff for the person who holds it; the payoff goes to future strangers and agents. That is a collective-action problem, and it needs process: make documentation part of incident closure, part of deploy checklists, and part of onboarding. The knowledge that gets written is the knowledge with a forcing function attached.
Edge cases and pitfalls
- Do not document secrets or credentials; reference the secret store path instead.
- Beware the "comprehensive wiki" trap: a thousand unreviewed pages are less useful than a hundred trusted ones.
- Agents take docs literally, so write preconditions and scope explicitly ("applies to the payments cluster only").
- Include negative knowledge: "we tried X and it failed because Y" saves the agent from re-discovering dead ends.
- Assign owners per section; ownerless docs are the first to rot.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_VwMpcvTZPIQT1D28US0JiQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.