# "unusual API activity" alert: investigation steps

## TL;DR
An API anomaly alert usually means a key is being used in a way it wasnt before. Identify the caller, compare the last hour to the last 30 days, and look for the two bad patterns: brand-new endpoints being touched and error rates spiking. Rotate the key if anything smells off; a rotated key costs nothing compared to a leaked one.

```text
"unusual API activity" alert: investigation steps
```

## Use this when
- your monitoring fires an unusual API activity alert
- API usage graphs show a spike you cant explain
- a partner reports odd calls coming from your integration
- you need a repeatable triage for API anomalies

## Not for this skill when
- you are designing API rate limits (that is prevention)
- the anomaly is in webhook deliveries, not API calls
- you need to debug a broken integration (that is engineering, not security)
- the API in question is a third party's, and you are the caller (check with them)

## Steps
1. Identify the caller. Pull the key ID, the service or user it belongs to, the source IPs, and the user agent strings. A key you issued to a backend service suddenly calling from a residential IP is already a strong signal.

2. Build the 30-day baseline. Pull the last 30 days of calls for that key, grouped by endpoint and hour. The alert fired because something deviated; this tells you what "normal" looks like so you can name the deviation.

3. Check for new endpoints. List what the key touched in the alert window and compare it against the baseline. Admin, export, and user-listing endpoints that the key never touched before are the story:
```spl
index=api key_id=[key id]
| stats count by endpoint
| sort - count
```
Expected: the same handful of endpoints as the baseline. New rows near the top mean someone is exploring.

4. Check error rates. A jump in 401s or 403s suggests someone probing permissions they dont have. A jump in 429s suggests scraping or brute force. A jump in 5xx errors might just be your bug, so check the deploy log too.

5. Check the key's story. When was it created, when was it last rotated, how broad are its scopes, and has it shown up anywhere it shouldnt (code repos, logs, chat). Old, broad, never-rotated keys are the ones that leak.

6. Decide and act. Rotate the key, tighten its scopes to what the service actually needs, block abusive source IPs, and tell the key owner what changed and why. Then watch the new key for a week to make sure the anomaly doesnt follow it.

### Variant: investigating GraphQL anomalies
With GraphQL there are no distinct endpoints, so look at query depth, field selection, and result sizes instead. A sudden jump in deeply nested queries or bulk field selection is the GraphQL version of "new endpoints."

### Variant: webhook abuse patterns
If the anomaly is inbound webhooks, check for replayed payloads, signature validation failures, and event types the sender never sent before. A spike in invalid signatures means someone found your webhook URL.

### Variant: third-party integration keys behaving oddly
When the key belongs to a vendor's integration, loop the vendor in early. Ask what changed on their side before you rotate; some anomalies are their new feature, not your breach. Rotate anyway if you cant get a clean answer fast.

## Why this happens
API keys are long-lived credentials that get copied into code, logs, dashboards, and chat. When one leaks, the usage pattern changes before anyone notices the leak itself. The anomaly alert is usually the first sign, which is why the investigation starts with "what changed" rather than "who leaked it."

## Edge cases and pitfalls
- Legitimate traffic spikes look like anomalies. Product launches, batch jobs, and migrations all do this. Check the calendar before you accuse anyone.
- Shared keys make caller identity fuzzy. If five services share one key, you cant tell which one is misbehaving. Move to per-service keys after the incident.
- Dont revoke a production key at 2am without a replacement ready. Have the new key issued and distributed before you kill the old one, or you trade a security incident for an outage.
- Attackers throttle themselves to dodge anomaly detection. If the alert was borderline, look at longer windows; slow abuse hides in hourly averages.
- Log the key ID, not the key value, in your investigation notes. The value is a secret; treat it like one.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_syOpuKaCZEvw602T2TLaug
