flyway migrate hit the command timeout but the schema change kept running in the background after the agent gave up
A recovery playbook for when an agent's flyway migrate is killed by a command timeout but the schema change may have finished on the server: how to read flyway_schema_history, verify against the real schema, and repair or resume without double-applying. Use when flyway timed out and the migration state is unknown. Not for migrations that failed with a SQL error.
TL;DR
The timeout killed the agent's client session, not necessarily the migration: the database may have finished, rolled back, or be halfway through. Read flywayschemahistory for the latest rows, check the actual schema to see whether the change landed, then use flyway repair followed by flyway migrate. Never re-run migrate blind: check state first or you get checksum and duplicate-object failures.
The query
flyway migrate hit the command timeout but the schema change kept running in the background after the agent gave upUse this when
- flyway migrate was killed by a command or harness timeout
- The agent gave up but the schema change may have completed on the server
- flywayschemahistory and the agent's memory disagree
- You need to decide whether it is safe to re-run
Not for
- Migrations that failed with an actual SQL error (fix the SQL first)
- flyway validate checksum failures on a healthy history (that is a file-editing problem)
- Baseline or repair on a database flyway has never seen (different operation)
Steps
1. Read the migration history
SELECT installed_rank, version, description, type, script, checksum,
installed_by, installed_on, execution_time, success
FROM flyway_schema_history
ORDER BY installed_rank DESC
LIMIT 5;Expected output: the latest rows with their success flags. success = 1 means flyway recorded it as applied; success = 0 (or a missing row) means it never completed from flyway's point of view.
2. Check whether the schema change actually landed
Query the real schema for the objects the timed-out migration was supposed to create or alter - informationschema.tables, informationschema.columns, pg_indexes, whatever the migration script touched.
Expected output: a yes-or-no answer for each object: the change is either in the database or it is not. This is the ground truth; the history table is only flyway's opinion.
3. Reconcile the two and pick the path
- History says success and the change is present: nothing to do. The migration finished; the agent just did not see it.
- History says failed or missing, and the change is absent: the migration rolled back. Go to step 4.
- History says failed or missing, but the change IS present: the DDL committed but flyway never recorded it. Go to step 5.
Expected output: one of the three cases identified, no guessing.
4. Repair, then migrate
flyway repair
flyway migraterepair removes the failed history row; migrate then applies the migration cleanly from scratch.
Expected output: repair reports the failed row removed, and migrate applies the migration with success = 1.
5. The change landed but was never recorded: repair, then verify
flyway repair
flyway validaterepair clears the failed row. Do NOT run migrate here: the schema change already exists, and re-running would fail on duplicate objects. If validate complains about a missing migration, the correct fix is flyway repair plus a documented manual mark, not a blind re-run.
Expected output: validate passes and the history table matches the real schema.
6. Prevent the next one
Raise the command timeout for long migrations instead of letting the agent's default kill them, and give the agent a pre-flight step: read flywayschemahistory before every migrate so a resume starts from facts.
Expected output: long migrations get a timeout budget that matches their runtime, and the agent checks state before running.
Variant phrasings
flyway migrate timed out but kept running
Same playbook. Steps 1 and 2 first: the timeout tells you nothing about what the database did.
flyway schema history says failed but table exists
Step 5. The DDL committed; repair the history and do not re-run the migration.
safe to re-run flyway migrate after timeout
Only after steps 1 through 3 say the change is absent. If the change is present, re-running is exactly what corrupts the state.
Why it happens
A command timeout kills the client connection, but the database server keeps working on whatever it was doing until it notices the client is gone. Postgres DDL is transactional, so a killed session usually rolls the migration back - but "usually" is doing a lot of work: statements that already committed, non-transactional migrations, and history-table writes that race the kill all produce the split-brain state where the schema and the history disagree. The agent sees a timeout and assumes nothing happened; the database may disagree.
Edge cases
- Non-transactional migrations (Postgres CREATE INDEX CONCURRENTLY and friends): these cannot roll back. A timeout mid-migration leaves a partial state that step 2 must map object by object.
- Two agents racing the same migrate: the loser gets a lock or duplicate-key failure, not a timeout. Serialize migration runs.
- flyway repair on a shared database affects every environment reading that history table: coordinate before running it on a shared schema.
- The agent's harness timeout and flyway's own connect timeout are different knobs: set both, and make the harness budget larger than flyway's.
- If validate keeps failing after repair, someone edited an already-applied migration file: that is a checksum problem, not a timeout problem.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_2021YGV2s4qYl5WBfbw0aA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.