# how to rotate API keys without downtime

## TL;DR

Never swap a key in one step. Create the new key while the old one still works, deploy the new key to every consumer, verify traffic has moved by watching usage logs, then revoke the old key. Overlap is the whole trick; everything else is just being thorough about finding every consumer.

```text
how to rotate API keys without downtime
```

## Use this when

- A key is due for scheduled rotation
- A key may have been exposed and needs replacing fast
- Someone with key access leaves the team
- You are rotating database credentials or service tokens

## Not for this skill when

- The secret is in git history (rotate first, then use the cleanup skill)
- You need to choose where secrets live (different skill)
- It is a human password, not a machine credential

## Steps

### 1. Inventory every consumer of the key

List everywhere the key is used: services, CI jobs, cron scripts, integrations, dev setups, dashboards. The consumer you forget is the outage.

```bash
grep -rn "[key-name-or-identifier]" --include="*.yml" --include="*.yaml" --include="*.json" --include="*.env*" ~/workspace/deploy-configs/ | grep -v node_modules | cut -d: -f1 | sort -u
```

Expected: a file list of every config referencing the key. Cross-check against the provider's usage dashboard; configs you cant find by grep (CI secret stores, vendor dashboards) are the ones that bite.

### 2. Create the new key alongside the old one

Generate the replacement while the old key stays active; most providers support multiple active keys per account.

```bash
aws iam create-access-key --user-name [service-user]
```

Expected: a new access key pair returned. Store it in your secret manager immediately (never in chat, never in a ticket); note both key IDs and which is old vs new.

### 3. Deploy the new key to every consumer

Roll the new key out through your normal config deployment in waves, keeping the old key working throughout.

```bash
vault kv put secret/myapp api_key [paste the new key value here]
```

Expected: the new value is in the secret store. Trigger redeploys or restarts so each consumer picks it up; verify each wave before starting the next.

### 4. Verify traffic moved to the new key

Watch the provider's usage logs for the old key ID going quiet and the new one taking over. Dont proceed until the old key shows zero traffic for a full cycle (a day for daily jobs, an hour for hot paths).

In the CloudTrail console, look up events for the old access key ID via Event history filtered by access key.

Expected: no recent events for the old key, or only events older than your quiet window. Any consumer still on the old key shows up here; fix it before revoking.

### 5. Revoke the old key and confirm

Only now, disable then delete the old key. Disabling first (rather than deleting) gives you a one-step rollback if something was missed; after a quiet period, delete it.

```bash
aws iam update-access-key --user-name [service-user] --access-key-id [old-key-id] --status Inactive
```

Expected: the old key is inactive. Watch error rates for a bit; if something breaks, reactivate, find the missed consumer, and repeat step 4. After a clean quiet period, delete the key.

### Variant: emergency rotation after possible exposure

Same steps, compressed: create and deploy the new key immediately, skip the long quiet window (watch for 15 minutes of clean traffic instead of a day), then revoke. Speed matters more than elegance, but dont skip the consumer inventory or you trade a leak for an outage.

### Variant: rotating database passwords without downtime

Databases often allow two valid passwords during rotation (RDS and most managed DBs support this). Set the new password as secondary, roll the app, promote, remove the old. For self-hosted Postgres, do it at the role level with a brief overlap window.

### Variant: rotating keys an employee had access to

Rotate everything they could reach, not just what you think they used. People copy keys into personal scripts and local env files; the inventory in step 1 has to assume the worst. Do it promptly; delayed rotation after offboarding is how "former employee" becomes "incident."

## Why this happens

Keys are shared secrets with no expiry, so they accumulate: copied into dashboards, baked into scripts, saved in password managers of people who left two years ago. Rotation bounds that sprawl, but a hard swap breaks every consumer you forgot about. Overlap turns rotation from a flag day into a non-event.

## Edge cases and pitfalls

- The forgotten consumer: there is always one; the usage-log check in step 4 is the safety net, not the inventory.
- Keys baked into container images or mobile apps: these cant be updated by redeploy; you need an app release or a server-side kill switch.
- Rate limits on key creation: some providers cap active keys at two; plan the overlap so you never need three.
- CI caches: a cached build artifact with the old key keeps working until the cache expires; bust caches as part of the rollout.
- Webhook secrets at vendors: rotating means updating the secret on both sides; do the vendor side first so deliveries keep verifying.
- Deleting instead of disabling: deletion is irreversible on most providers; disable, wait, then delete.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_1-Zqo8RG8sLCfPkOBSK4Jg
