how to build an approval gate for agent actions
A step-by-step skill for adding human or policy approval before AI agents take consequential actions: action classification, approval UX, timeouts, and audit records. Use when an agent or engineer is asked to put a human in the loop, require sign-off for risky agent operations, or design tiered autonomy. Triggers: 'agent approval', 'human in the loop', 'approve agent actions'. Not for: general workflow approvals, CI/CD manual gates, or access-request systems.
TL;DR
Classify agent actions by blast radius, auto-approve the safe ones, and require explicit approval for everything irreversible or external, with a clear summary of what will happen and a default-deny on timeout. An approval gate that approves everything by default, or that nobody reads, is decoration.
The query
how to build an approval gate for agent actionsUse this when
- Your agent can send, delete, publish, pay, or change production
- You want tiered autonomy: full auto for safe actions, human sign-off for risky ones
- Compliance requires a human decision on certain operations
- You are tired of the agent asking about trivial things and want the line drawn well
Not for
- Generic business workflow approvals
- CI/CD manual deployment gates
- Identity access request flows
Steps
- Classify actions into tiers by blast radius: read-only (auto-approve), reversible within your systems (auto-approve with logging), irreversible or external (require approval). Write the tiers down with examples per tool.
Expected output: a tier table covering every tool the agent can call.
- Design the approval request: what action, with what arguments, on whose behalf, what happens if approved, and how to undo it. The approver decides from this card alone; if they need to dig, the card is wrong.
Expected output: a sample approval card a non-expert can understand in ten seconds.
- Enforce the gate at the dispatch layer, not in the agent's prompt. The agent requests; a separate service approves or denies. Prompt-only gates are suggestions the agent can talk itself out of.
Expected output: a test where the agent is told to skip approval still gets blocked.
- Set timeouts and defaults: approval requests expire, and expiry means deny, never approve. Batch related requests so approvers are not spammed into clicking yes blindly.
Expected output: an unanswered request is denied and logged after the timeout.
- Log every decision: who approved, what they saw, when, and the resulting action. Approvals are audit events, tie them into the tool-call log.
Expected output: any past action traces back to its approval record.
- Tune the line: track approval volume, deny rate, and time-to-decision. If approvers rubber-stamp, your tiers are wrong: move more into auto-approve-with-logging or make the risky ones genuinely rare.
Expected output: a weekly report showing the gate is used, not bypassed.
Variant phrasings
"human in the loop for AI agents"
Same pattern. The key design choice is which loop: per-action approval vs per-plan approval. Per-plan scales better; per-action is safer for novel tasks.
"agent action policy engine"
The automated version of the gate: rules approve or deny without a human for the middle tier. Same tiers, same dispatch-layer enforcement, rules instead of clicks.
"approve agent tool calls in Slack"
A common UX for the gate: the approval card goes where the team already is. Keep the card complete; a bare "approve?" button trains people to click yes.
Why this happens
Agents are fast and literal; humans are slow and contextual. Without a gate, the agent's speed becomes your incident speed. With a bad gate, humans become a rubber stamp and you get the worst of both. The tier system exists to spend human attention where it matters.
Edge cases and pitfalls
- Approval fatigue is the top failure mode; fewer, better approval requests beat more.
- Standing approvals ("always allow this") need expiry and review, or they quietly become blanket permission.
- The agent must not be able to approve its own actions or rephrase a denied action to dodge the gate; match on action semantics, not wording.
- Emergency break-glass needs its own audited path; do not make people bypass the gate off the books.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_Ny3g2CuplzF4pRsABSyxcQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.