# Intel agent hung on a 503 retry loop from the transcript API

## TL;DR
A transcript API 503 loop hangs the agent the same way a news API loop does: unbounded retries against an outage. Bound the retries, open a breaker, and continue the briefing without transcripts for this cycle. The transcript jobs queue for the next run; the briefing ships on time without them.

## The error
```text
(agent hung)
intel agent hung on 503 retry loop from transcript api; retries unbounded, run never finished
```

## When this helps
- an agent hangs on transcript API 503s
- transcript outages stall briefings
- building breakers for transcription APIs
- deferring transcript work

## When it doesn't
- the 503 is brief; short backoff suffices
- transcripts are the whole briefing; then waiting may be justified
- the API is retired; migrate, do not break around it

## Works with
python 3.8+ with requests and json. Any transcript API.

## Steps
### 1. Bound transcript API retries
```python
import time, requests
s = requests.Session()
for i in range(5):
    r = s.get("https://api.YOUR-transcript-provider/v1/jobs", timeout=20)
    print("try", i + 1, r.status_code)
    if r.status_code != 503:
        break
    time.sleep(10 * (i + 1))
else:
    print("breaker open for transcript api")
```
Expected: Five tries then stop. The loop's else branch fires only on total failure.

### 2. Queue transcript jobs for the next run
```python
import json
pending = ["job1", "job2"]
open("pending_transcripts.json", "w").write(json.dumps(pending))
print("transcript jobs deferred, not lost")
```
Expected: A deferred queue. The outage delays transcripts; it does not delete them.

### 3. Brief without transcripts this cycle
```python
import json
print("briefing proceeds on filings and news; transcript section marked pending")
open("coverage_note.txt", "w").write("transcripts deferred: provider 503")
```
Expected: A briefing without the transcript section. The gap is disclosed, not hidden.

### 4. Re-probe the API on the next scheduled run
```python
import json
print("next run: one probe call before resuming transcript pulls")
print("success closes the breaker; failure keeps it open")
```
Expected: A probe policy. Recovery is automatic and bounded.

## Other ways people phrase this
### transcript api 503 retry loop
Bound retries, defer jobs, brief without. Probe next run.

### agent hung transcript api outage
The hang is the agent's design. Breakers fix the design.

### 503 transcript jobs stuck
Deferred queues preserve the work across the outage.

## Why it happens
Transcript APIs have outages like any service, and unbounded retries turn a minutes-long outage into a hung run. The breaker pattern bounds the damage, and because transcripts are one input among many, the briefing degrades gracefully without them.

## Edge cases
- Transcript outages during earnings week hurt most; the fallback matters then.
- Do not let deferred jobs pile up across many runs; cap the queue.
- A provider with chronic 503s needs replacing, not more breakers.
- Log outage windows; they explain briefing gaps later.
- If the provider's status page shows an incident, skip the probe and keep the breaker open until it clears.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_nlokBUd-6IV6R9rXTTgqXQ
