## TL;DR
A PVC that never binds with a provisioner error means the external provisioner cannot talk to the storage backend, not that the PVC is wrong. Check the PVC events for the provisioner name, confirm that provisioner pod is running, read its logs for the backend error, and verify its RBAC and backend credentials. The PVC spec is usually fine; the provisioner-to-backend path is broken.

## Error / query
```text
how to debug storage class provisioner failures
```

## Use this skill when
- PVCs stay Pending with `failed to provision volume` events
- Dynamic provisioning worked before and stopped
- The external-provisioner sidecar logs backend errors
- A new storage class never successfully provisions

## Not for this skill when
- You bind PVCs to pre-created PVs manually (no provisioner involved)
- PVCs are Pending for node affinity or topology (scheduling)
- Using local static provisioner with pre-provisioned disks

## Steps

### Step 1: Read the PVC events for the provisioner name
```bash
kubectl describe pvc [pvc-name] -n [namespace] | grep -A10 Events
```
Expected: `failed to provision volume with StorageClass "[sc]": ...` naming the provisioner (e.g. `ebs.csi.aws.com`). The text after the colon is the backend error; that is your real lead.

### Step 2: Confirm the provisioner pod is healthy
```bash
kubectl get pods -n [provisioner-namespace] -l app=[provisioner-app]
```
Expected: provisioner pods Running. If they are CrashLooping or Pending, fix that first; no provisioner means no provisioning, and the PVC events will just repeat.

### Step 3: Read the provisioner's backend error
```bash
kubectl logs -n [provisioner-namespace] deploy/[provisioner-deploy] -c [provisioner-container] --tail=50 | grep -i "error\|fail" | tail -10
```
Expected: the storage-backend error (auth rejected, quota exceeded, wrong region/zone, API throttling). This is the actionable message; PVC events only summarize it.

### Step 4: Check provisioner RBAC and backend credentials
```bash
kubectl get clusterrole [provisioner-role] -o yaml | grep -c "persistentvolumes"
kubectl get secret [backend-secret] -n [provisioner-namespace]
```
Expected: the ClusterRole covers persistentvolumes/persistentvolumeclaims/storageclasses, and the backend secret exists with current credentials. Expired cloud credentials are the most common silent cause.

## Variant phrasings

### "failed to provision volume with storageclass"
The PVC event text. Steps 1-3 trace it to the provisioner and the backend error.

### "pvc pending provisioner error"
Same flow. The provisioner pod logs (step 3) have the answer the PVC events summarize.

## Why it happens
Dynamic provisioning is a chain: PVC to StorageClass to provisioner pod to storage backend API. The PVC only reports the last error it saw, which hides whether the break is at the provisioner (down, RBAC) or the backend (auth, quota, region). Reading down the chain in order finds the actual break fast.

## Edge cases and pitfalls
- Zone/region mismatches: the provisioner may create the volume in a zone with no nodes that can attach it; check `allowedTopologies` on the StorageClass.
- API throttling from the cloud provider looks like intermittent failures; back off and check quotas, not the provisioner config.
- Multiple provisioners with similar names (in-tree vs CSI) cause confusion; verify the StorageClass `provisioner:` field matches the running provisioner.
- Deleting a stuck PVC does not clean up a half-created backend volume; check the storage console for orphans after repeated failures.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_Vjklm73b8fyA5Ix5XMgCxw
