## TL;DR
Make the agent show its work before doing anything: kubectl diff or --dry-run=server for Kubernetes, terraform plan for infra, helm diff for charts. A human or policy approves the plan, then a separate step applies exactly what was approved. Plan and apply are never the same command.

## Error / query
```text
dry-run and plan-first patterns for infra agents
```

## Use this skill when
- an agent proposes Kubernetes manifest changes
- an agent manages Terraform or CloudFormation
- you are designing the agent's change workflow
- a past incident came from an agent applying without review

## Not for this skill when
- the agent is read-only and never mutates anything
- changes go through a full GitOps pipeline already (the pipeline is the plan step)
- you need emergency break-glass without review (document that path separately)

## Steps

### Step 1: Have the agent render the intended change as a diff first
```bash
kubectl diff -f manifests/ -n prod
```
Expected: a unified diff of exactly what would change on each object. If the agent cannot produce a clean diff, it does not understand the change well enough to apply it.

### Step 2: Use server-side dry-run for admission-aware validation
```bash
kubectl apply --dry-run=server -f manifests/ -n prod
```
Expected: "created/configured (server dry run)" or validation errors. Server dry-run runs admission webhooks and policy checks, catching what client dry-run misses; treat a dry-run failure as a hard stop.

### Step 3: For Terraform, require a saved plan artifact
```bash
terraform plan -out=tfplan -detailed-exitcode; echo "exit=$?"
terraform show -json tfplan | jq ".resource_changes | length"
```
Expected: exit 2 means changes pending, and the JSON shows the count and type of resource changes. Store the tfplan file as the review artifact; the apply step must use that exact file, not a fresh plan.

### Step 4: Gate the apply on explicit approval of the plan
```bash
terraform apply tfplan
kubectl apply -f manifests/ -n prod
```
Expected: apply succeeds with no surprises beyond the approved diff. This step runs only after a human or policy approved step 1-3's output. Never let the agent chain plan and apply in one go.

### Step 5: Verify the applied state matches the plan
```bash
kubectl diff -f manifests/ -n prod
terraform plan -detailed-exitcode; echo "exit=$?"
```
Expected: kubectl diff shows no changes and terraform plan exits 0. Post-apply drift means something else moved, or the apply did not do what the plan said; investigate before calling it done.

## Variant phrasings

### "agent terraform workflow safe pattern"
Plan to a file, human reviews the JSON summary, apply uses the saved file, then a fresh plan confirms zero drift. The saved plan is the contract.

### "preview helm changes before agent applies"
Use helm template piped into kubectl diff, or helm diff upgrade, as the step 1 equivalent; same rule that nothing applies without a reviewed preview.

### "can an agent just run kubectl apply"
Only against dev namespaces with tight RBAC and even then prefer the diff habit. Against prod, never without the plan-first gate.

## Why it happens
Apply-without-preview fails because the agent's mental model of current state is often stale or wrong: someone else changed the object, the base manifest drifted, or the agent misread a selector. The diff forces reality into the loop before anything mutates, and the approval step puts a second brain on the change. Most agent-caused outages are "it applied what it thought was true", which plan-first prevents by construction.

## Edge cases and pitfalls
- kubectl diff needs the same kubeconfig and namespace as the apply; a diff against the wrong cluster is worse than no diff.
- Server dry-run can still miss mutating webhooks that behave differently on real admission; keep the approval gate regardless.
- Terraform plans expire when state or variables change; re-plan if anything moved between approval and apply.
- Agents love to "fix" a failed dry-run by loosening the change; require re-review of the new diff, not silent iteration.
- Large plans hide destructive changes in noise; fail the review automatically if the plan contains any delete or replace of stateful resources.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_5yWtx2FgcS5dTrY0XHZrXA
