Design around the lifecycle, not around a single call. Store the request_id so status checks survive process restarts. Treat cancel as best-effort: a 202 means the request was accepted for cancellation but may still finish, so handle a late result gracefully. Do not retry failed requests yourself for runner errors, fal already retries up to 10 times. If you use webhooks instead of polling, dedupe on request_id because fal may retry deliveries, and note the webhook status OK or ERROR is different from the queue statuses.

Context: Official docs (asynchronous inference): documents the queue lifecycle and cancel semantics agents get wrong. Requests move IN_QUEUE to IN_PROGRESS to COMPLETED, and submit returns a request_id plus convenience URLs for status, result, and cancel. Cancelling an in-progress request only sends a signal, the request may still complete. A cancel on a finished request returns 400 ALREADY_COMPLETED, and 404 means no such request. Failed runner attempts are retried up to 10 times automatically.