cost guardrails for agents on BigQuery/Snowflake
Caps agent spend on BigQuery and Snowflake with three layers: per-query warehouse limits, per-session budgets tracked in your app, and a kill switch on the agent role. Use when agents run ad hoc warehouse queries, when surprise bills are a risk, or when you need per-agent spend attribution. Do not use as a replacement for query result caching, for storage cost control, or for human query budgets which need different tooling.
TL;DR
Cap spend in three layers: per-query limits enforced by the warehouse itself, per-session budgets tracked in your application, and a kill switch on the agent's role. Agents have no sense of money, so the budget has to live outside the agent. It works because the warehouse-level cap is the one that holds when everything else fails: even a confused agent cannot spend what the database refuses to bill.
cost guardrails for agents on BigQuery/SnowflakeUse this when
- Agents run ad hoc warehouse queries on your bill
- A surprise five-figure invoice is a realistic fear
- You need per-agent or per-session spend attribution
- Leadership wants a hard number on maximum agent data cost
Not for
- Storage cost control, which is a separate problem
- Human analyst budgets, which need different tooling
- Replacing query caching, which reduces spend instead of capping it
Steps
- Set a per-query byte cap on BigQuery:
from google.cloud import bigquery
job_config = bigquery.QueryJobConfig(
maximum_bytes_billed=10_000_000_000, # 10 GB max per query
labels={"agent_id": "agent-7", "session": "2026-10-04-001"},
dry_run=False,
)Expected output: any query that would scan more than 10 GB fails immediately with a billing-limit error instead of running up charges. The agent gets an error it can react to, not a bill you discover later.
- Dry-run expensive-looking queries first to estimate cost:
dry_config = bigquery.QueryJobConfig(dry_run=True)
dry_job = client.query(sql, job_config=dry_config)
print(f"estimated bytes: {dry_job.total_bytes_processed:,}")Expected output: the byte estimate before execution. If it exceeds the session budget, the agent must narrow the query instead of running it.
- On Snowflake, use resource monitors plus statement timeouts:
CREATE RESOURCE MONITOR agent_monitor WITH CREDIT_QUOTA = 50
TRIGGERS ON 80 PERCENT DO NOTIFY ON 100 PERCENT DO SUSPEND;
ALTER WAREHOUSE agent_wh SET RESOURCE_MONITOR = agent_monitor;
ALTER USER agent_svc SET STATEMENT_TIMEOUT_IN_SECONDS = 120;Expected output: the warehouse suspends itself at the credit quota, and no single query runs longer than two minutes. Both are enforced by Snowflake, not by your code.
- Track a per-session budget in your application and stop at the limit:
session_spend = ledger.get_spend(session_id) # from the cost ledger
if session_spend + estimated_cost > SESSION_BUDGET_USD:
raise BudgetExceeded(f"session {session_id}: ${session_spend:.2f} of ${SESSION_BUDGET_USD:.2f} used")Expected output: the agent gets a clean budget error with numbers, which it can report to the user instead of silently burning money.
- Keep a kill switch on the agent role for emergencies:
-- one statement, immediate effect, no side effects on humans
SUSPEND WAREHOUSE agent_wh;
-- or: REVOKE ROLE agent_reader FROM USER agent_svc;Expected output: all agent spend stops within seconds. Document who may pull it and require a post-incident note when they do.
Variant phrasings
limit BigQuery cost per query agent
maximumbytesbilled is the single most effective control on BigQuery. Set it on every agent query path, not just the main one.
Snowflake budget cap AI agent
Resource monitors cap the warehouse, statement timeouts cap the query. You need both, because a cheap-but-endless query and an expensive-but-fast query fail in different ways.
prevent LLM expensive queries
Prevention is layered: estimate before running, cap per query, budget per session, kill switch for emergencies. Any single layer has a hole; together they hold.
Why it happens
Warehouses price curiosity: every exploratory query scans real bytes or burns real credits, and agents are extremely curious, retrying with slight variations until something works. A human feels the hesitation before running a query that might cost $40; an agent feels nothing. The guardrails externalize the hesitation into mechanisms that do not depend on the agent's judgment, which is exactly the part you cannot rely on.
Edge cases
- Dry-run estimates miss slot-time costs and some engine overheads; treat estimates as floors, not ceilings.
- Cached results cost nothing but can mask the true price of a query pattern; do not set budgets from cached runs alone.
- On flat-rate or capacity pricing the dollars work differently, but the byte and credit caps still prevent runaway consumption of shared capacity.
- An agent can split one expensive question into many cheap queries that sum past the per-query cap; the session budget in step 4 is what catches this.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstlCQBgcZgbtZNeJdJBdkzQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.