## TL;DR

Django records every fully applied migration in the `django_migrations` table, so a timed-out run leaves an exact record of what landed. Run `showmigrations --plan` to list unapplied migrations in order, check whether the interrupted migration's DDL actually landed in the database, and re-run `migrate` to resume. Django skips already-applied migrations, so the safe move is to resume, not to guess.

## The query

```text
django migrate timed out on a big AddField in production and now the agent cant tell which migrations actually applied
```

## Use this when

- `python manage.py migrate` was killed by a timeout or the connection dropped mid-run in production.
- The agent needs to know which migrations applied before deciding whether to retry.
- The timed-out migration was a heavy DDL operation like a big AddField.

## Not for

- Migration conflicts from two branches (that is a merge conflict, not a timeout resume).
- Squashed migrations reporting "no such column" or missing-table errors.
- A migration that fails deterministically every time it runs (that is a broken migration, not an interrupted one).

## Steps

### Step 1: List what Django thinks is applied

```bash
python manage.py showmigrations --plan | grep -v "\[X\]"
```

Expected output: the migrations Django still considers unapplied, in dependency order. Everything marked `[X]` applied cleanly before the timeout.

### Step 2: Check whether the interrupted migration's DDL actually landed

```sql
SELECT column_name
FROM information_schema.columns
WHERE table_name = '[table]' AND column_name = '[column]';
```

Expected output: if the column exists, the ALTER committed before the timeout killed the client. If it does not exist, the DDL was rolled back or never ran. This check decides the next step, so do not skip it.

### Step 3a: If the column exists but the migration is unapplied, mark it applied

```bash
python manage.py migrate [app] [migration] --fake
```

Expected output: Django records the migration as applied without running its DDL. Use this only when step 2 proved the DDL landed. Faking a migration whose DDL did not run corrupts the schema.

### Step 3b: If the column does not exist, just re-run migrate

```bash
python manage.py migrate
```

Expected output: Django applies the pending migrations in order, skipping everything already recorded in `django_migrations`. The big AddField runs again from scratch.

### Step 4: Guard the retry against another timeout

```bash
python manage.py migrate --database default
```

Expected output: same as step 3b, but run it inside a session with a raised statement timeout or from a host with a stable connection. A big AddField on a large production table can take far longer than the default tool timeout, so raise the timeout before retrying or the loop repeats.

### Step 5: Verify the plan is empty

```bash
python manage.py showmigrations --plan | grep -v "\[X\]" | wc -l
```

Expected output: 0. Every migration is marked applied. Run the app's smoke checks against the migrated database before closing the incident.

## Variant phrasings

### django migrate was killed halfway - how to tell which migrations applied
`showmigrations --plan` plus the `django_migrations` table give the exact applied set. Steps 1-2 are the reconciliation.

### agent re-ran django migrate after a timeout and it tried to re-apply everything
It did not re-apply. Django skips applied migrations. If it looks like a re-apply, the first run never recorded them, which means step 2's DDL check matters even more.

### big AddField migration keeps timing out in production
The ALTER itself is the problem, not the resume logic. Consider a concurrent index-style approach or a multi-step migration (add nullable, backfill, then set NOT NULL) instead of one giant DDL.

## Why it happens

Django wraps each migration in a transaction where the database supports transactional DDL (PostgreSQL does, MySQL does not fully). A timeout kills the client, not necessarily the DDL: on PostgreSQL the in-flight ALTER either committed or rolled back, and `django_migrations` only records fully completed migrations. The ambiguity the agent feels is real but resolvable: the table says what completed, and the information schema says what the DDL did.

## Edge cases

- On MySQL, DDL is not transactional, so a killed AddField can leave a half-altered table. Check the table structure carefully before re-running.
- Never `--fake` a migration unless step 2 proved its DDL landed. A faked-but-missing column breaks every later migration that references it.
- If two deploys run `migrate` concurrently, both can attempt the same migration. Serialize production migrates with a deploy lock.
- `showmigrations --plan` shows the plan for the default database. For multi-database setups, pass `--database` for each one.
- A migration that times out because of a lock (not slowness) will time out again. Check `pg_stat_activity` for blockers before retrying.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_Fmp3EcoynUXpOUnxq52toA
