## TL;DR
Start chaos engineering with experiments so small they feel boring: kill one pod in staging, add latency to one dependency, fill one disk to 80 percent. Each experiment states a hypothesis (`we believe the system does X when Y happens`), runs with a big red abort button, and teaches you one thing. Boring first experiments build the safety habits that make bigger ones possible.

## Error / query
```text
first chaos experiment ideas for a small team
```

## Use this skill when
- Starting chaos engineering with a small team
- Picking first experiments with minimal risk
- Proving value before buying tooling
- The team is nervous about breaking things

## Not for this skill when
- Setting up chaos tooling (installation guide)
- Running large game days (different format)
- Chaos at scale across many services

## Steps

### Step 1: Pick the smallest meaningful blast radius
```text
Candidate first experiments:
1. Kill one pod of a stateless service in staging.
2. Add 500ms latency to a non-critical dependency.
3. Fill a staging disk to 80 percent.
4. Restart one node in a dev cluster.
Rule: staging first, one fault at a time, business hours only.
```
Expected: a shortlist everyone agrees is safe. If anyone is uncomfortable, shrink the blast radius further; confidence comes from uneventful runs.

### Step 2: Write the hypothesis before running
```text
Template: "We believe [system] will [expected behavior] when [fault].
We will abort if [condition]. Steady state is [metric + value]."
Example: "We believe checkout stays under 2s p99 when one pod dies.
Abort if p99 exceeds 5s for 2 minutes. Steady state: p99 800ms."
```
Expected: a falsifiable prediction plus abort criteria. The hypothesis is the point; the fault is just the probe. No hypothesis, no experiment.

### Step 3: Run with monitoring and a hand on the abort
```bash
# watch the steady-state metric live during the experiment
# abort command ready: stop the fault injection immediately
```
Expected: you watch the metric, not the fault tool. If the abort condition hits, stop the experiment; an aborted experiment that taught you the limit is a success, not a failure.

### Step 4: Write down what you learned and pick the next one
```text
One paragraph: hypothesis, what happened, what surprised you,
what you will fix or test next. File it where the team can find it.
```
Expected: a growing log of validated (or invalidated) assumptions. The log compounds: after ten experiments you have a map of what the system actually does under stress, which is worth more than any single fix.

## Variant phrasings

### "chaos engineering small team start"
Steps 1-2. Tiny blast radius, written hypothesis, staging first.

### "safe first chaos experiments"
The candidate list in step 1. Boring is the goal for experiment one.

## Why it happens
Teams avoid chaos engineering because the imagined version is `randomly break production`. The real version starts with trivially safe experiments that validate assumptions cheaply. Small teams especially benefit: they cannot afford a dedicated chaos team, but they also cannot afford to discover their failover is broken during a real incident.

## Edge cases and pitfalls
- Never run the first experiments in production; earn production with a streak of clean staging runs.
- Experiments without abort criteria are just outages you scheduled; the abort condition is non-negotiable.
- Third-party dependencies in the blast radius need their own consideration; your staging experiment should not DDoS a vendor's sandbox.
- If an experiment reveals a real weakness, fix it before running bigger experiments; stacking unknowns teaches nothing.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_7zdRraT364gyFn9_hRnbjA
