## TL;DR
The timeout killed the agent's client session, not necessarily the migration: the database may have finished, rolled back, or be halfway through. Read flyway_schema_history for the latest rows, check the actual schema to see whether the change landed, then use flyway repair followed by flyway migrate. Never re-run migrate blind: check state first or you get checksum and duplicate-object failures.

## The query

```text
flyway migrate hit the command timeout but the schema change kept running in the background after the agent gave up
```

## Use this when

- flyway migrate was killed by a command or harness timeout
- The agent gave up but the schema change may have completed on the server
- flyway_schema_history and the agent's memory disagree
- You need to decide whether it is safe to re-run

## Not for

- Migrations that failed with an actual SQL error (fix the SQL first)
- flyway validate checksum failures on a healthy history (that is a file-editing problem)
- Baseline or repair on a database flyway has never seen (different operation)

## Steps

### 1. Read the migration history

```sql
SELECT installed_rank, version, description, type, script, checksum,
       installed_by, installed_on, execution_time, success
FROM flyway_schema_history
ORDER BY installed_rank DESC
LIMIT 5;
```

Expected output: the latest rows with their success flags. success = 1 means flyway recorded it as applied; success = 0 (or a missing row) means it never completed from flyway's point of view.

### 2. Check whether the schema change actually landed

Query the real schema for the objects the timed-out migration was supposed to create or alter - information_schema.tables, information_schema.columns, pg_indexes, whatever the migration script touched.

Expected output: a yes-or-no answer for each object: the change is either in the database or it is not. This is the ground truth; the history table is only flyway's opinion.

### 3. Reconcile the two and pick the path

- History says success and the change is present: nothing to do. The migration finished; the agent just did not see it.
- History says failed or missing, and the change is absent: the migration rolled back. Go to step 4.
- History says failed or missing, but the change IS present: the DDL committed but flyway never recorded it. Go to step 5.

Expected output: one of the three cases identified, no guessing.

### 4. Repair, then migrate

```bash
flyway repair
flyway migrate
```

repair removes the failed history row; migrate then applies the migration cleanly from scratch.

Expected output: repair reports the failed row removed, and migrate applies the migration with success = 1.

### 5. The change landed but was never recorded: repair, then verify

```bash
flyway repair
flyway validate
```

repair clears the failed row. Do NOT run migrate here: the schema change already exists, and re-running would fail on duplicate objects. If validate complains about a missing migration, the correct fix is flyway repair plus a documented manual mark, not a blind re-run.

Expected output: validate passes and the history table matches the real schema.

### 6. Prevent the next one

Raise the command timeout for long migrations instead of letting the agent's default kill them, and give the agent a pre-flight step: read flyway_schema_history before every migrate so a resume starts from facts.

Expected output: long migrations get a timeout budget that matches their runtime, and the agent checks state before running.

## Variant phrasings

### flyway migrate timed out but kept running

Same playbook. Steps 1 and 2 first: the timeout tells you nothing about what the database did.

### flyway schema history says failed but table exists

Step 5. The DDL committed; repair the history and do not re-run the migration.

### safe to re-run flyway migrate after timeout

Only after steps 1 through 3 say the change is absent. If the change is present, re-running is exactly what corrupts the state.

## Why it happens

A command timeout kills the client connection, but the database server keeps working on whatever it was doing until it notices the client is gone. Postgres DDL is transactional, so a killed session usually rolls the migration back - but "usually" is doing a lot of work: statements that already committed, non-transactional migrations, and history-table writes that race the kill all produce the split-brain state where the schema and the history disagree. The agent sees a timeout and assumes nothing happened; the database may disagree.

## Edge cases

- Non-transactional migrations (Postgres CREATE INDEX CONCURRENTLY and friends): these cannot roll back. A timeout mid-migration leaves a partial state that step 2 must map object by object.
- Two agents racing the same migrate: the loser gets a lock or duplicate-key failure, not a timeout. Serialize migration runs.
- flyway repair on a shared database affects every environment reading that history table: coordinate before running it on a shared schema.
- The agent's harness timeout and flyway's own connect timeout are different knobs: set both, and make the harness budget larger than flyway's.
- If validate keeps failing after repair, someone edited an already-applied migration file: that is a checksum problem, not a timeout problem.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_2021YGV2s4qYl5WBfbw0aA
