## TL;DR
"Notify retry canceled" means Alertmanager gave up delivering to your webhook after retries: the endpoint is down, too slow, or returning errors. The problem is almost always the receiving end, not Alertmanager. Check the webhook endpoint's health and response codes, then look at what changed on that side. Alertmanager's retry logic is generous; cancellation means persistent failure.

## The query
```text
Alertmanager webhook receiver "notify retry canceled": how to debug
```

## Use this when
- Webhook notifications never arrive
- Alertmanager logs show retry canceled
- After changing the webhook endpoint
- Notifications worked before and stopped

## Not for when
- Email or PagerDuty receiver issues
- Alert routing (alerts never reach the receiver)
- Alert rule evaluation problems

## Steps

### Step 1: Check the webhook endpoint health
Verify the receiving endpoint is up and accepting POSTs. A down or redeployed receiver is the most common cause; Alertmanager retries for a while, then cancels, exactly matching the symptom.
Expected output: the endpoint health confirmed, or the outage found.

### Step 2: Inspect response codes
Check what the endpoint returns: 4xx means the payload or auth is wrong (Alertmanager will keep failing); 5xx means the receiver is broken; timeouts mean it is overloaded or network-partitioned. The code determines the fix.
Expected output: the response class directing the fix at payload, auth, or receiver health.

### Step 3: Validate the payload format
If the endpoint returns 400s, it dislikes the payload: check the expected JSON shape against what Alertmanager sends. Receiver upgrades that change the expected format break previously working webhooks.
Expected output: payload shape aligned with the receiver's expectations.

### Step 4: Check authentication
Verify webhook secrets, tokens, or signatures. Rotated credentials on either side produce persistent 401s that Alertmanager retries until cancellation.
Expected output: valid auth on both sides.

### Step 5: Test delivery independently
Send a test alert payload to the endpoint manually and confirm it processes. Then trigger a real test alert through Alertmanager. The manual test separates endpoint problems from Alertmanager problems.
Expected output: end-to-end delivery proven working.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_IXTzv_3ivDXYgAg7yFg7bA
