how to debug RBAC forbidden errors after upgrades
Debugs RBAC forbidden errors that appear after cluster upgrades. Use when kubectl or workloads suddenly get forbidden errors post-upgrade, service accounts lose permissions, or aggregated roles stop working. Covers auth can-i checks, ClusterRole bindings, and version-removed RBAC rules. Not for authentication failures, OIDC misconfiguration, or initial RBAC setup.
TL;DR
Post-upgrade forbidden errors mean the upgrade changed what a role is allowed to do, not who you are. Check what the caller can do with kubectl auth can-i, diff the ClusterRoles against the pre-upgrade versions, and look for RBAC rules that referenced removed or renamed API groups. Re-apply the current version's default roles rather than hand-patching old ones.
Error / query
how to debug RBAC forbidden errors after upgradesUse this skill when
kubectlcommands that worked before the upgrade now returnforbidden- Service accounts lose permissions after a cluster upgrade
- Controllers or operators break with RBAC errors post-upgrade
- Error text names a resource and verb, e.g.
cannot list resource "x"
Not for this skill when
- The error is
unauthorized(authentication, not authorization) - RBAC never worked (initial setup problem)
- OIDC or webhook auth is misconfigured (identity problem)
Steps
Step 1: Reproduce and read the exact denial
kubectl auth can-i [verb] [resource] --as=[user-or-sa] -n [namespace]
# example: kubectl auth can-i list deployments --as=system:serviceaccount:[ns]:[sa] -n [ns]Expected: no confirms the denial, and the original error names the exact resource, verb, and API group. The API group matters: upgrades move resources between groups (e.g. extensions to apps).
Step 2: Find which bindings grant that subject
kubectl get clusterrolebindings,rolebindings -A -o json | jq -r '.items[] | select(.subjects[]? | .name=="[sa-name]") | .metadata.name'Expected: the bindings attaching roles to the subject. If a binding is missing that used to exist, the upgrade's manifests may have replaced it; if it exists, the role's rules are the suspect.
Step 3: Diff the role rules against the API groups that moved
kubectl get clusterrole [role-name] -o yaml | grep -A10 "apiGroups"Expected: rules referencing current API groups. Classic post-upgrade break: a rule allows extensions/v1beta1 deployments but the cluster now serves apps/v1; the rule silently matches nothing. Update rules to the new group/version.
Step 4: Re-apply the current default roles instead of patching
kubectl auth can-i [verb] [resource] --as=[user-or-sa] -n [namespace]Expected: after fixing the role (correct API group, correct resource name, correct verb), can-i returns yes. Prefer re-applying the role definition shipped with the new cluster version over editing the old one in place; old roles accumulate stale groups.
Variant phrasings
"forbidden after kubernetes upgrade"
Steps 1-3. API group moves between versions are the usual cause.
"serviceaccount forbidden cannot list"
The service account's role references a stale API group. Update the rule to the group the cluster actually serves now.
Why it happens
Kubernetes deprecates and removes API versions on a schedule, and RBAC rules match on exact group/version/resource strings. A role written for extensions/v1beta1 keeps existing after the upgrade but matches zero requests, so everything it granted silently stops working. Upgrades can also replace default ClusterRoles and bindings with new definitions.
Edge cases and pitfalls
kubectl auth can-i --listshows the full effective permission set; use it when the denial message is vague.- Aggregated ClusterRoles (with aggregationRule) pick up rules from labeled roles; an upgrade can change the labels the aggregator selects.
- Impersonation (
--as) needs its own permission; ifcan-i --asitself is forbidden, test from the real identity instead. - Do not grant
*on*to make the error go away; scope the fix to the exact resource and verb in the denial.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstL8HtIu-iM7YRdqlzWI8Rw