building a shared library of verified data fixes
Builds a shared library of verified data fixes: each entry records the exact symptom, the solution that worked, how it was verified, and an owner, in a searchable repo the agent queries before writing new SQL. Use when the same data errors recur, when incident knowledge lives in chat threads, or when agents keep reinventing solutions to solved problems. Do not use for unverified guesses, for fixes too specific to ever recur, or as a replacement for fixing root causes.
TL;DR
Store every fix as a small record: the symptom in exact error text, the verified solution that worked, how you proved it, and an owner, in a searchable repo the agent queries before writing new SQL. A fix is not verified until it ran against real data and someone confirmed the output. It works because the same data errors recur constantly, and each recorded incident makes the next one cheaper while stopping the agent from inventing solutions to solved problems.
building a shared library of verified data fixesUse this when
- The same data errors hit the team every few months
- Incident knowledge lives in chat threads nobody can find
- Agents keep reinventing solutions to solved problems
- Onboarding new team members to the team's scar tissue
Not for
- Unverified guesses that happened to work once
- Fixes so specific they will never recur
- Replacing the work of fixing the actual root cause
Steps
- Define the record format and keep it small enough to actually fill in:
# fixes/2026-10-04-null-customer-id.yaml
symptom: "nulls appeared in customer_id on analytics.orders"
error_text: "dbt test not_null_orders_customer_id failed: 3 violations"
diagnosis: "left join to refunds fanned out on multi-refund orders"
solution_sql: "pre-aggregate refunds per order_id before joining (see diff)"
verified_how: "reran dbt test green; row counts match source; reviewer: Priya"
owner: data-platform
applies_when: "any model joining orders to a one-to-many table"
last_confirmed: 2026-10-04Expected output: a file a future debugger can read in thirty seconds and apply in five minutes.
- Set the verification bar and enforce it at intake:
VERIFIED means all three:
[ ] reproduced on real data (not just reasoned about)
[ ] the fix ran and the check went green
[ ] a second person confirmed the output
UNVERIFIED entries go in fixes/drafts/, never in the searched library.Expected output: a clear line between folklore and verified fixes. The drafts folder is where guesses wait; the library is where answers live.
- Make the agent query the library before writing new SQL:
matches = fix_library.search(symptom=agent_error_text, top_k=3)
if matches:
agent_context += "Known fixes for similar errors:\n" + format(matches)Expected output: the agent sees verified fixes for similar symptoms before generating its own solution. Most recurring errors get solved from the library on the first try.
- Review the library on a cadence and deprecate what rotted:
-- quarterly: fixes not confirmed in 180 days get flagged
SELECT symptom, owner, last_confirmed FROM ops.fix_library
WHERE last_confirmed < CURRENT_DATE() - 180;Expected output: a list of stale entries for owners to re-confirm or retire. A fix for a schema that no longer exists is worse than no fix; it is a confident wrong answer.
- Link fixes to runbooks so the library connects to operations:
# in the runbook for the orders pipeline
known_issues:
- fix: fixes/2026-10-04-null-customer-id.yaml
check: "dbt test not_null_orders_customer_id"Expected output: on-call engineers and agents find the fix from the alert, not from a separate search. The fix library and the runbook reference each other.
Variant phrasings
data fixes knowledge base
The knowledge is the verified records. A wiki page describing the fix without the verification fields is documentation, not a library.
verified SQL fix library
Verification is the differentiator: reproduced, fixed, confirmed by a second person. Everything else is a snippet collection.
incident fix repository for agents
The agent integration in step 3 is what makes it a repository for agents rather than a repository agents happen to be able to read.
Why it happens
Data teams solve the same five problems forever: fan-out joins, timezone bugs, null handling, late data, and duplicate loads. Without a library, each occurrence costs a fresh investigation, and agents, which have no memory of your incidents, pay that cost every single time. The library converts incidents into assets: the pain of debugging once buys cheap resolution forever. The verification bar is what keeps it trustworthy, because an agent acting on a wrong fix with high confidence is worse than an agent with no fix at all.
Edge cases
- Overly specific fixes never match future symptoms; write the applies_when field to capture the general pattern, not just the incident.
- Ownership without maintenance: assign library review to a rotation, not to whoever filed the fix, or entries rot when people move teams.
- Search quality determines usefulness; tag entries with error text verbatim so exact-match queries hit, and with pattern tags for fuzzy ones.
- Fixes that paper over root causes deserve a linked ticket to the real fix; the library entry should say what it does not fix.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_QLhm9-tyEXuVm0IxHZRNkQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.