## TL;DR
Throw away the partial output, rebuild from clean into a staging directory, and only swap it into place when sphinx-build exits 0. The crash is only half the problem; the other half is the next build or deploy picking up half-written pages. Clean, rebuild, gate the swap on success.

```text
docs agent failed: crashed mid sphinx build, left partial output
```

## Steps
1. Confirm the crash and the mess: check the agent log for the sphinx-build traceback or kill signal, and list the output dir. Expected: missing pages or a build timestamp older than the crash.
2. Delete the partial output: `rm -rf build/html` (or your configured output dir). Expected: the directory is gone; nothing half-written remains to confuse the next build.
3. Rebuild from clean into a staging dir: `sphinx-build -b html docs/source build/staging`. Expected: exit code 0 and a "build finished" line; any new traceback names the real culprit (often one bad page or extension, find it with -v).
4. Only publish on success: move the staging output into the live location after a verified exit 0, never during the build. Expected: the served site is either the full new build or the full old one, never a mix.
5. Harden the agent loop: on a nonzero exit it discards the output and retries once, then stops and reports the traceback instead of looping forever. Expected: a repeat crash produces a clear error, not another half-built site.

## Use this when
- sphinx-build crashes or is killed mid-run
- the output dir holds a half-built site after an agent run
- deploys sometimes serve a mix of old and new pages

## Not for this skill when
- the build fails deterministically on a specific page (fix the page or extension, then rebuild)
- the build succeeds but output looks wrong (content bug)
- the agent crashes outside the build step (different failure)

## Variant phrasings
- sphinx build died halfway, partial html left
- agent crashed during sphinx-build
- half-built sphinx output after crash

## Why it happens
sphinx-build writes pages incrementally, so a crash (extension exception, out-of-memory kill, agent timeout) freezes the output mid-write. The next run or deploy then sees a directory that looks like a finished build but is missing pages, and happily serves it.

## Edge cases
- Out-of-memory kills leave no traceback: check system logs for the memory killer before blaming an extension.
- A fresh-environment rebuild is slower but rules out stale-cache weirdness when the crash cause is unclear.
- Parallel builds can crash on extensions that are not thread-safe: retry single-threaded to isolate.
- If the crash is an agent timeout, not a real crash, raise the timeout instead of just retrying.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_OMOqenyrearQWOdS43PiYQ
