## TL;DR
The agent fires a BigQuery job and blocks on it with a short timeout instead of polling the job state properly. Fix it by submitting the job, then polling `job.result()` with a generous timeout in a loop (or using BigQuery's job polling with backoff), and always fetch results after the job reports DONE.

```text
data agent timed out waiting for bigquery job: polling pattern
```

## Use this when
- A data agent times out waiting for a BigQuery job
- Long queries die at the agent's HTTP timeout, not BigQuery's
- The job actually succeeded but the agent never collected results

## Not for this skill when
- The query itself errors (thats the SQL)
- Quota blocks the job (thats capacity)
- The query is slow (thats tuning, a different job)

## Steps

1. Submit without blocking, then poll the job state:

```python
job = client.query("SELECT ...")  # returns immediately
print(job.job_id, job.state)
```
Expected output: the job id and state (RUNNING). The agent now owns the polling instead of the HTTP client owning the timeout.

2. Wait with a timeout that matches the workload, and handle it:

```python
try:
    rows = job.result(timeout=600)
except TimeoutError:
    print(f"job {job.job_id} still running; will re-check")
```
Expected output: either rows, or a controlled timeout where the job keeps running server-side. The job is not lost on timeout; only the wait gave up.

3. On timeout, re-attach to the same job instead of resubmitting:

```python
job = client.get_job(job.job_id)
rows = job.result(timeout=600)
```
Expected output: results from the original job. Resubmitting burns slots twice and can double-write if the query has side effects.

4. For fire-and-forget pipelines, poll with backoff in the orchestrator:

```python
import time
while job.state != "DONE":
    time.sleep(10)
    job.reload()
```
Expected output: the loop exits when the job finishes. Cap the loop with a max wait and alert rather than looping forever.

## Variant phrasings

### agent reports the query failed but BigQuery shows success
The wait timed out, not the query. The results are sitting in the job; re-attach (step 3).

### timeouts only on the first run of the day
Cold slots and cold caches make the first query slow. The polling pattern absorbs this; a fixed short timeout doesnt.

## Why it happens
BigQuery jobs are asynchronous by design: submit returns fast, execution takes as long as it takes. Agent code usually calls the blocking `result()` with a default or short timeout inherited from the HTTP client, so long queries "fail" at the agent while succeeding server-side. The polling pattern respects the async model instead of fighting it.

## Edge cases
- `job.result()` with no timeout blocks forever; always pass one in agent code.
- Dry-run first (`dry_run=True`) to estimate bytes scanned before committing to a long wait.
- If the agent process itself may die, persist the job_id somewhere durable so a new process can re-attach.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_BWINVBO9lsaNneTGgTZ3mg
