migration checkpoint file was written to /tmp and the container restarted - the agent cant tell what already migrated
Fixes migration state lost when a checkpoint file in /tmp vanished on container restart. Use when the agent can't tell what already migrated because its progress file is gone. Key trigger: /tmp checkpoint missing after restart; rebuild truth from the migration tool's state table and move checkpoints somewhere durable.
TL;DR: Treat the /tmp checkpoint as gone - in containers it always is after a restart. Rebuild the truth from the migration tool's state table in the database, diff it against the migration files on disk, and re-run only what's pending. Then move the checkpoint somewhere that survives restarts: a small table in the database itself or a file on a persistent volume.
The problem as reported:
migration checkpoint file was written to /tmp and the container restarted - the agent cant tell what already migrated- Confirm the file is really gone by listing its path. Expected: no such file. Don't spend time trying to recover it - it isn't coming back.
- Query the migration tool's state table for the applied set (
alembic current,flyway info,npx prisma migrate status,python manage.py showmigrations, orrails db:migrate:status). Expected: the definitive list of applied migrations. - Diff the applied set against the migration files on disk to get the pending list. Expected: a concrete list of migrations still to run.
- Re-run the migrate command. Expected: only pending migrations apply; the tool finishes with the database up to date.
- Replace the /tmp checkpoint with a durable one: either a migration_runs table in the database (run id, step, status, finished time) updated after each step, or a file on a mounted persistent volume. Expected: after a container restart, the agent reads the checkpoint and resumes instead of starting over.
- Test the fix: restart the container and have the agent resume mid-plan. Expected: it picks up at the recorded step with no rework.
Use this when
- A progress or checkpoint file was lost to a container restart or sandbox recycle.
- The agent can't tell what already migrated because its notes are gone.
- You need a checkpoint design that survives restarts.
Not for this skill when
- The checkpoint file still exists - just read it.
- The migration tool already tracks applied state durably (most do) - then you never needed the file in the first place; query the state table directly.
Variant phrasings
- migration progress file deleted after container restart
- /tmp checkpoint lost how to resume migration
- container restarted mid migration lost state
- checkpoint file gone can't tell what migrated
Why it happens
Container filesystems are ephemeral: /tmp is wiped on restart by design. Anything the next run needs - including "what already migrated" - has to live in the database or on a persistent volume.
Edge cases
- The checkpoint was ahead of reality (it said "done" for a step that actually failed) - the database state table wins every tie-break.
- Two agent instances writing the same checkpoint file - use one writer or a database row with a unique run id.
- A volume that looks persistent but is actually per-pod - verify the mount survives a real restart before trusting it.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_hXxOLQNaIA6yaIgxSLUAmw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.