## TL;DR
Do not point a new SRE agent at production on day one. Stand up a staging cluster that mirrors prod topology, throw scripted incidents at it, score the agent on detection, diagnosis, and safe remediation, and only promote it after it passes consistently. Staging is where the agent learns your environment cannot hurt anyone.

## Error / query
```text
how to test an SRE agent in staging first
```

## Use this skill when
- you are onboarding a new SRE agent or a new agent version
- the agent is getting new capabilities (write access, new tools)
- you need acceptance criteria before a prod rollout
- a staging environment exists but is not used for agent testing

## Not for this skill when
- the agent is read-only (lighter validation suffices)
- there is no staging environment at all (build a minimal one first)
- you are load-testing the agent harness itself

## Steps

### Step 1: Build a staging cluster that mirrors prod topology
```bash
kind create cluster --name agent-staging --config staging-cluster.yaml
kubectl --context kind-agent-staging get nodes
```
Expected: a running cluster with the same namespace layout and core addons as prod. It does not need prod scale; it needs prod shape, so the agent's commands and assumptions transfer.

### Step 2: Deploy a canary app with known failure modes
```bash
kubectl --context kind-agent-staging apply -f staging-apps/
kubectl --context kind-agent-staging -n demo rollout status deployment/demo-api
```
Expected: the demo app healthy. Include the failure modes you care about: crashlooping pods, OOM kills, bad config pushes, expired certs. Each one is a test case with a known right answer.

### Step 3: Script the incident scenarios, do not hand-wave them
```bash
kubectl --context kind-agent-staging -n demo set resources deployment/demo-api --limits=memory=64Mi
kubectl --context kind-agent-staging -n demo get pods -w
```
Expected: pods start OOMKilling within a minute or two. Scripted faults are reproducible; "just poke at it" is not a test. Keep a scenario library in git next to the staging config.

### Step 4: Run the agent against each scenario and score it
```bash
./run-agent-eval.sh --scenario oomkill --context kind-agent-staging | tee /tmp/eval-oomkill.log
grep -E "detected|diagnosed|remediated|escalated" /tmp/eval-oomkill.log | tail -10
```
Expected: the eval log shows the agent's actions scored on four axes: did it detect the issue, diagnose the root cause correctly, remediate safely (or escalate instead of guessing), and stay within its guardrails. An agent that "fixes" things by deleting them fails.

### Step 5: Promote only on a consistent pass record
```bash
ls /tmp/eval-*.log | wc -l
grep -l "GUARDRAIL_VIOLATION" /tmp/eval-*.log || echo "no violations"
```
Expected: a set of eval logs with zero guardrail violations across at least three consecutive full runs. One lucky pass is not a signal; consistency is. New agent versions re-run the whole suite before they touch prod.

## Variant phrasings

### "acceptance test for ops ai agent"
The four-axis score (detect, diagnose, remediate-or-escalate, stay in guardrails) across the scripted scenario library, with a documented pass threshold.

### "chaos engineering for agents"
Same idea with nastier faults: network partitions, clock skew, cascading failures. Start with single faults; combined faults are the advanced class.

### "agent did fine in staging but broke prod"
Then staging diverged from prod somewhere: compare topology, data shape, permissions, and traffic patterns, and add the missing realism to the scenario library.

## Why it happens
Agents behave differently in each environment because their actions depend on what they observe, and staging that "kind of" looks like prod produces confidence that does not transfer. The failure modes that bite are always the ones staging did not replicate: the weird CRD, the legacy namespace, the permission that exists in prod but not staging. Mirroring topology and scripting real faults closes most of that gap before prod pays for the rest.

## Edge cases and pitfalls
- Staging data must never be a copy of prod PII; synthesize or scrub it, even for agent testing.
- Agents can learn staging-specific shortcuts (hardcoded names, relaxed policies); randomize names and keep staging policies as strict as prod.
- Eval scoring needs a human in the loop at first; fully automated scoring of agent reasoning comes later, after you trust the rubric.
- Do not let staging rot; a staging cluster that drifts from prod for six months is a false sense of security.
- Time-box the staging phase per capability; perfect staging confidence is impossible, so define "good enough to canary in prod" explicitly.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_A8_P0QsMFUX8GsJzFqpMrQ
