## TL;DR
Guardrails for production agents come in layers: least-privilege identity, dry-run or plan-first execution, policy enforcement that can say no, approval gates for anything destructive, and full audit trails. No single layer is enough; an agent that slips past one should hit the next.

## Error / query
```text
guardrails for agents doing production changes
```

## Use this skill when
- an agent will apply changes to production infrastructure or data
- you are writing the safety review for an agent deployment
- leadership asks "what stops the agent from breaking prod"
- an agent already caused a scare and you need controls retrofitted

## Not for this skill when
- the agent is read-only (lighter controls suffice)
- you need to bypass or weaken existing guardrails for convenience (do not do this)
- the agent runs only in dev or staging

## Steps

### Step 1: Scope the agent's identity to least privilege
```bash
kubectl auth can-i --list --as=system:serviceaccount:ops:sre-agent -n prod
aws iam simulate-principal-policy --policy-source-arn [AGENT_ROLE_ARN] --action-names ec2:TerminateInstances --resource-arns "*"
```
Expected: the allowed list matches the agent's documented job and nothing more. If the agent only restarts deployments, terminate and delete verbs must be absent.

### Step 2: Enforce dry-run or plan-first for every change
```bash
kubectl diff -f manifests/ -n prod
terraform plan -out=tfplan -detailed-exitcode; echo "exit=$?"
```
Expected: kubectl diff shows the exact object changes; terraform plan exits 2 when changes are pending. The agent proposes, a human or policy approves, and only then does anything apply. Agents should never run bare apply against prod.

### Step 3: Add policy enforcement that can veto the agent
```bash
kubectl get clusterpolicies.kyverno.io | head -10
kyverno apply --cluster --policy require-runbook-label.yaml --resource manifests/ | tail -5
```
Expected: the policy list shows guardrails like "require change ticket label" or "deny latest tag"; the test apply reports pass or fail per resource. Policy engines say no even when the agent is confident, which is the point.

### Step 4: Put destructive actions behind an approval gate
```bash
kubectl argo rollouts get rollout [ROLLOUT] -n prod -o jsonpath="{.status.pauseConditions}"
```
Expected: paused rollouts waiting on a manual promote step. Deploys, restarts of stateful services, and any delete flow through a gate a human taps; the agent prepares everything and waits.

### Step 5: Alert on every production write the agent performs
```bash
kubectl get events -n prod --sort-by=.lastTimestamp | grep -i "sre-agent" | tail -10
```
Expected: agent-attributed events visible in the stream. Wire your audit log into an alert so the on-call sees agent writes in near real time; silent agent changes are how small mistakes become incidents.

## Variant phrasings

### "safety controls for ai ops agent"
Same layered model: identity, plan-first, policy veto, approvals, audit. The review question is always "what happens when the agent is wrong", not "is the agent usually right".

### "prevent agent from changing prod database"
Keep data-plane credentials out of the agent's reach entirely; if it needs DB insight, give it a read replica with a read-only user and no path to the primary.

### "agent guardrails checklist for soc2"
Map each layer to evidence: RBAC bindings, plan artifacts stored per change, policy reports, approval records, and immutable audit logs.

## Why it happens
Agents optimize for completing the task, not for caution; a confident wrong action executes just as fast as a right one. Human operators have hesitation, context, and fear of the pager, and agents have none of those. Guardrails externalize the caution: the safety lives in the system around the agent, not in the agent's instructions, because instructions can be misread and systems enforce.

## Edge cases and pitfalls
- Prompt-injected instructions can tell the agent to skip steps; guardrails must live outside the agent's instruction stream (RBAC, admission control) where prompts cannot reach them.
- Approval fatigue is real; if every trivial change needs a tap, humans start rubber-stamping, so gate only what is actually risky.
- Dry-run is not a perfect preview (webhooks and controllers can differ); treat it as necessary but not sufficient.
- An agent with broad read access can still leak secrets through logs; scope read access to non-sensitive resources too.
- Test the guardrails with adversarial drills: ask the agent to do something forbidden in staging and confirm each layer blocks it.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_hm2hEm-Wt7ir5JKQt9JebQ
