## TL;DR

Start by breaking one non-critical thing in staging on a schedule, watch what breaks, and fix the surprises before they happen in prod. Chaos engineering for a small team is just deliberate practice with guardrails. You dont need a platform; you need a hypothesis, a blast radius, and a revert button.

## Error / query

```text
chaos engineering basics for small teams
```

Youve heard chaos engineering is for companies with 500 engineers. You want the version that works with five.

## Use this skill when

- the team has never deliberately tested failure
- you want resilience practice without buying a chaos platform
- staging exists but nobody breaks things in it on purpose
- leadership asks what chaos engineering would look like at your size

## Not for this skill when

- you need a game day template for incident process practice (a different skill)
- you need production chaos tooling evaluation (a buying decision)
- youre in a regulated industry needing formal DR proof (heavier process required)

## Steps

### 1. Pick one small experiment with a clear blast radius

Write the hypothesis, the blast radius, and the abort condition before touching anything. If you cant write the abort condition, the experiment is too big.

```bash
cat > /tmp/chaos-experiment.md <<'EOF'
# Experiment 1: kill one web pod in staging

Hypothesis: traffic shifts to the surviving pods with no 5xx spike
Blast radius: staging only, one pod, 5 minutes
Abort: if error rate exceeds 1% for 60s, stop immediately
EOF
cat /tmp/chaos-experiment.md
```

Expected: a written experiment with a hypothesis, a bounded blast radius, and an abort condition. No experiment runs without all three.

### 2. Run it and watch the dashboard

Delete the pod, then watch the error rate for the full window. The watching is the experiment; the deletion is just the trigger.

```bash
kubectl delete pod -l app=[app-name] -n staging
echo "pod deleted; watching error rate for 5 minutes"
sleep 300
echo "window complete; check the dashboard for 5xx spikes"
```

Expected: the pod terminates, surviving pods take the traffic, and the error rate stays flat. A spike is also a result: it found a gap.

### 3. Record what you learned

Log the date, the experiment, and the result. The log is what turns one-off experiments into institutional knowledge.

```bash
echo "- $(date -u +%Y-%m-%d): killed one staging pod, no 5xx spike, failover clean" >> ./chaos-log.md
cat ./chaos-log.md
```

Expected: an append-only log of experiments and outcomes. Future experiments build on past results instead of repeating them.

## Variant phrasings

### getting started with chaos engineering

Same fix: one small staging experiment with a hypothesis and an abort condition is the start. Everything else is scaling up.

### chaos monkey for small teams

Same fix: you dont need the monkey; a cron job that kills a staging pod weekly plus the log in step 3 is the small-team version.

### resilience testing basics

Same fix: hypothesize, break, observe, record. Thats resilience testing with or without the chaos label.

## Why it happens

Systems fail in ways nobody predicted because nobody tested failure on purpose. Staging exists precisely for this, but most teams use it only to verify that the happy path works. Deliberate failure in staging finds the missing health check, the single point of failure, and the bad timeout before production finds them for you.

## Edge cases and pitfalls

- No staging environment: run in prod off-hours with a tight blast radius and an abort plan. Smaller blast radius, same method.
- Experiment caused a real outage: thats data. Write it up blamelessly; the experiment did its job, just louder than planned.
- Team is scared to break things: start with read-only experiments like latency injection, not kills. Build the muscle gradually.
- Regulated industry: get the blast radius approved in writing first. The method is the same; the paperwork isnt optional.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_TvnCwgUjTOamtri8CFkTbw
