how to keep agent state in files that survive restarts
Shows how to keep agent state in files that survive restarts: state file patterns, atomic writes, watermarks, and recovery on boot. Use it when an agent must resume work after a crash or a VM restart. Triggered by questions about persistent agent state, surviving restarts, or state files vs databases. Not for database-backed state or for secrets storage.
TL;DR
Keep state in small JSON files, write them atomically (write temp, then rename), and record a watermark of how far the work got after every unit completes. On boot, read the watermark and resume from there. Agents that keep state only in memory lose everything on restart; agents with atomic state files resume as if nothing happened.
how to keep agent state in files that survive restartsUse this when
- An agent runs long pipelines that outlive a single session
- The VM or container can restart mid-run
- You need to resume a queue, a crawl, or a campaign after interruption
- You are debugging "it forgot where it was" failures
- You are choosing between files and a database for state
Not for this skill when
- You need concurrent writers with transactions (use a database)
- The state includes secrets (use a vault, never a state file)
- The state is large (gigabytes) or needs querying (use a database)
- The process never restarts in practice (memory is fine, do not over-engineer)
Steps
- Decide what state must survive, and keep it small. Queue position, completed item ids, per-item results, config watermarks. If it can be recomputed cheaply, do not store it; every stored byte is a byte that can go stale.
survives: [queue position, completed ids, results]
recomputed: [everything else]Expected output: an explicit list. Success check: the state file stays under a megabyte for your workload.
- Write state files atomically, always. Write to a temp file in the same directory, then rename over the target; readers never see a half-written file, even if the writer dies mid-write. This one habit prevents the majority of state corruption.
python3 -c "import json,os; json.dump(state, open('state.json.tmp','w')); os.rename('state.json.tmp','state.json')"Expected output: state updates that survive kill -9. Success check: kill the writer mid-update ten times; the file is always valid JSON.
- Save progress after every unit of work, not at the end. Mark each queue item done (with its result) the moment it completes, before moving to the next. End-of-run saves turn every crash into a full rerun.
[ ] result persisted per item, immediately after completion
[ ] no "save everything at the end" step existsExpected output: per-item persistence. Success check: killing the run loses at most one item's work.
- On boot, load state and resume from the watermark, then verify. Read the state file, skip completed items, re-verify any item marked in-progress (it may have half-finished), and continue. Boot logic is part of the program, not an afterthought.
[ ] boot loads state and resumes without manual steps
[ ] in-progress items re-verified, not blindly skippedExpected output: unattended resume. Success check: restart mid-run and confirm zero duplicates and zero skipped items.
- Keep state files out of fragile locations. Not in /tmp (wiped on reboot), not only on an ephemeral disk, and backed up if the work matters. State that lives somewhere volatile is state that does not survive.
state lives in: [durable path], backed up: [yes/no + how]Expected output: a durable, known location. Success check: a full VM restart loses nothing.
Variant phrasings
Agent checkpointing best practices
Checkpointing phrasing. Steps 2 and 3: atomic writes plus per-unit saves are the whole technique.
How to resume a crashed agent run
Recovery phrasing. Step 4: load the watermark, re-verify in-progress items, continue; the boot path is the feature.
Files vs database for agent state
Storage phrasing. Files win for single-writer, small, simple state; databases win for concurrent writers and queries. Step 1's size test decides.
Watermark pattern for pipelines
Pattern phrasing. The watermark is "how far did we get", saved per unit (step 3), read on boot (step 4).
Why it happens
Long-running agents die: OOM kills, VM restarts, deploys, spot terminations. Memory-only state dies with the process, and end-of-run saves die with the crash, so the only state that reliably survives is state written durably and incrementally. Atomic writes matter because the crash does not schedule itself around your save points; the rename guarantees readers see the last complete state, never a torn one.
Edge cases / pitfalls
- JSON corruption from non-atomic writes is the classic failure. If you take one thing from this skill, take step 2.
- Clock skew can confuse "last updated" watermarks across machines. Prefer monotonic counters or item ids over timestamps for resume position.
- State files grow. Completed-id lists for million-item queues need compaction (a completed-up-to watermark plus an exception list); plan for it before the file gets slow.
- Do not store credentials in state files. Watermarks and ids only; secrets live in the vault and get re-resolved on boot.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstQmfI5nnL7rQPs59FyX-Pw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.