how to debug storage class provisioner failures
Debugs storage class provisioner failures when PVCs never get volumes. Use when PVCs sit Pending with provisioner errors, dynamic provisioning stopped working, or the external-provisioner pod logs errors. Covers provisioner identity, RBAC, and backend connectivity. Not for manually created PVs, node-local storage, or PVCs blocked by scheduling.
TL;DR
A PVC that never binds with a provisioner error means the external provisioner cannot talk to the storage backend, not that the PVC is wrong. Check the PVC events for the provisioner name, confirm that provisioner pod is running, read its logs for the backend error, and verify its RBAC and backend credentials. The PVC spec is usually fine; the provisioner-to-backend path is broken.
Error / query
how to debug storage class provisioner failuresUse this skill when
- PVCs stay Pending with
failed to provision volumeevents - Dynamic provisioning worked before and stopped
- The external-provisioner sidecar logs backend errors
- A new storage class never successfully provisions
Not for this skill when
- You bind PVCs to pre-created PVs manually (no provisioner involved)
- PVCs are Pending for node affinity or topology (scheduling)
- Using local static provisioner with pre-provisioned disks
Steps
Step 1: Read the PVC events for the provisioner name
kubectl describe pvc [pvc-name] -n [namespace] | grep -A10 EventsExpected: failed to provision volume with StorageClass "[sc]": ... naming the provisioner (e.g. ebs.csi.aws.com). The text after the colon is the backend error; that is your real lead.
Step 2: Confirm the provisioner pod is healthy
kubectl get pods -n [provisioner-namespace] -l app=[provisioner-app]Expected: provisioner pods Running. If they are CrashLooping or Pending, fix that first; no provisioner means no provisioning, and the PVC events will just repeat.
Step 3: Read the provisioner's backend error
kubectl logs -n [provisioner-namespace] deploy/[provisioner-deploy] -c [provisioner-container] --tail=50 | grep -i "error\|fail" | tail -10Expected: the storage-backend error (auth rejected, quota exceeded, wrong region/zone, API throttling). This is the actionable message; PVC events only summarize it.
Step 4: Check provisioner RBAC and backend credentials
kubectl get clusterrole [provisioner-role] -o yaml | grep -c "persistentvolumes"
kubectl get secret [backend-secret] -n [provisioner-namespace]Expected: the ClusterRole covers persistentvolumes/persistentvolumeclaims/storageclasses, and the backend secret exists with current credentials. Expired cloud credentials are the most common silent cause.
Variant phrasings
"failed to provision volume with storageclass"
The PVC event text. Steps 1-3 trace it to the provisioner and the backend error.
"pvc pending provisioner error"
Same flow. The provisioner pod logs (step 3) have the answer the PVC events summarize.
Why it happens
Dynamic provisioning is a chain: PVC to StorageClass to provisioner pod to storage backend API. The PVC only reports the last error it saw, which hides whether the break is at the provisioner (down, RBAC) or the backend (auth, quota, region). Reading down the chain in order finds the actual break fast.
Edge cases and pitfalls
- Zone/region mismatches: the provisioner may create the volume in a zone with no nodes that can attach it; check
allowedTopologieson the StorageClass. - API throttling from the cloud provider looks like intermittent failures; back off and check quotas, not the provisioner config.
- Multiple provisioners with similar names (in-tree vs CSI) cause confusion; verify the StorageClass
provisioner:field matches the running provisioner. - Deleting a stuck PVC does not clean up a half-created backend volume; check the storage console for orphans after repeated failures.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_Vjklm73b8fyA5Ix5XMgCxw