bigquery slot contention reservation explained
Explains BigQuery slot contention and reservations for data agents. Use when queries queue instead of running, when you need to understand on-demand vs reserved slots, or when deciding whether autoscaling reservations are worth it. Not for clustering vs partitioning decisions, for Snowflake warehouse issues, or for dbt model configuration.
TL;DR
BigQuery runs queries on slots (compute units). On-demand gives you a shared pool with burst limits; when your queries queue, you are out of slots. Reservations buy dedicated capacity, and the fix for chronic queuing is either fewer concurrent heavy queries or a reservation sized to your peak.
bigquery slot contention reservation explainedUse this when
- Queries sit queued instead of executing
- You are choosing on-demand vs reservations
- Slot usage spikes unpredictably
Not for this skill when
- You are deciding clustering vs partitioning
- Snowflake queries spill to remote disk
- You are configuring dbt models
Steps
- Confirm the symptom is slot contention, not query shape. In the BigQuery UI, queued queries show wait time before execution:
Look at INFORMATION_SCHEMA.JOBS: check total_slot_ms vs queued time.
High queued time with modest total_slot_ms = contention.
High total_slot_ms with no queue = the query itself is heavy.Expected output: the distinction. Do not buy slots to fix a query that scans terabytes it does not need.
- See who is eating the slots:
SELECT user_email, count(*) AS jobs, sum(total_slot_ms)/1000 AS slot_seconds
FROM `region-us`.INFORMATION_SCHEMA.JOBS
WHERE creation_time > TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 1 HOUR)
GROUP BY user_email ORDER BY slot_seconds DESC;Expected output: slot consumption by principal over the last hour. Usually one service account or one scheduled batch dominates.
- For bursty agent workloads, on-demand with burst is often fine, but know its limits:
On-demand: you share a pool, BigQuery bursts you above baseline briefly,
then throttles. Sustained heavy use queues.
Flat-rate/reservations: dedicated slots, no queuing up to capacity,
you pay whether you use them or not.Expected output: the decision framework. Agents firing ad-hoc queries in bursts fit on-demand; a fleet of agents querying all day fits reservations.
- Size a reservation from measured peak, not vibes:
SELECT TIMESTAMP_TRUNC(period_start, HOUR) AS hr,
max(avg_slots) AS peak_slots
FROM `region-us`.INFORMATION_SCHEMA.JOBS_TIMELINE
WHERE period_start > TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 7 DAY)
GROUP BY hr ORDER BY peak_slots DESC LIMIT 10;Expected output: your actual hourly slot peaks over the last week. Buy the reservation near the sustained peak, not the single highest spike.
- Isolate workloads so one hungry job cannot starve the rest:
Create separate reservations (or assign jobs to reservation assignments)
for: interactive/agent queries vs scheduled batch pipelines.
The batch can queue without paging anyone; the agents cannot.Expected output: noisy-neighbor protection. This is the operational fix that matters most once you have reservations.
Variant phrasings
bigquery queries queued waiting for slots
Slot contention (step 1). Reduce concurrency, lighten the queries, or buy reservations.
bigquery on demand vs flat rate
On-demand suits spiky, unpredictable use; reservations suit sustained high use. Measure with JOBS_TIMELINE before choosing (step 4).
bigquery reservation autoscaling
Autoscaling reservations grow within bounds you set, paying for burst only when it happens. Good middle ground for agent workloads with unpredictable peaks.
Why it happens
BigQuery's compute is slots, and on-demand projects share a finite pool per region. Your queries compete with everyone else's for burst capacity; sustained overuse gets throttled into a queue. Reservations convert the shared pool into dedicated capacity you control.
Edge cases
- Slot contention is per-region; moving a dataset to a less busy region is occasionally the real fix.
- Idle reservations still bill, set autoscale minimums low for dev projects.
INFORMATION_SCHEMAqueries themselves consume a little; sample rather than scanning a year.- Query queuing also happens per-user with fair scheduling; one user cannot always starve another even on shared capacity.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_mMJz1jAYemIJqWYHwRAhHA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.