## TL;DR
Post-upgrade `forbidden` errors mean the upgrade changed what a role is allowed to do, not who you are. Check what the caller can do with `kubectl auth can-i`, diff the ClusterRoles against the pre-upgrade versions, and look for RBAC rules that referenced removed or renamed API groups. Re-apply the current version's default roles rather than hand-patching old ones.

## Error / query
```text
how to debug RBAC forbidden errors after upgrades
```

## Use this skill when
- `kubectl` commands that worked before the upgrade now return `forbidden`
- Service accounts lose permissions after a cluster upgrade
- Controllers or operators break with RBAC errors post-upgrade
- Error text names a resource and verb, e.g. `cannot list resource "x"`

## Not for this skill when
- The error is `unauthorized` (authentication, not authorization)
- RBAC never worked (initial setup problem)
- OIDC or webhook auth is misconfigured (identity problem)

## Steps

### Step 1: Reproduce and read the exact denial
```bash
kubectl auth can-i [verb] [resource] --as=[user-or-sa] -n [namespace]
# example: kubectl auth can-i list deployments --as=system:serviceaccount:[ns]:[sa] -n [ns]
```
Expected: `no` confirms the denial, and the original error names the exact resource, verb, and API group. The API group matters: upgrades move resources between groups (e.g. `extensions` to `apps`).

### Step 2: Find which bindings grant that subject
```bash
kubectl get clusterrolebindings,rolebindings -A -o json | jq -r '.items[] | select(.subjects[]? | .name=="[sa-name]") | .metadata.name'
```
Expected: the bindings attaching roles to the subject. If a binding is missing that used to exist, the upgrade's manifests may have replaced it; if it exists, the role's rules are the suspect.

### Step 3: Diff the role rules against the API groups that moved
```bash
kubectl get clusterrole [role-name] -o yaml | grep -A10 "apiGroups"
```
Expected: rules referencing current API groups. Classic post-upgrade break: a rule allows `extensions/v1beta1` deployments but the cluster now serves `apps/v1`; the rule silently matches nothing. Update rules to the new group/version.

### Step 4: Re-apply the current default roles instead of patching
```bash
kubectl auth can-i [verb] [resource] --as=[user-or-sa] -n [namespace]
```
Expected: after fixing the role (correct API group, correct resource name, correct verb), `can-i` returns `yes`. Prefer re-applying the role definition shipped with the new cluster version over editing the old one in place; old roles accumulate stale groups.

## Variant phrasings

### "forbidden after kubernetes upgrade"
Steps 1-3. API group moves between versions are the usual cause.

### "serviceaccount forbidden cannot list"
The service account's role references a stale API group. Update the rule to the group the cluster actually serves now.

## Why it happens
Kubernetes deprecates and removes API versions on a schedule, and RBAC rules match on exact group/version/resource strings. A role written for `extensions/v1beta1` keeps existing after the upgrade but matches zero requests, so everything it granted silently stops working. Upgrades can also replace default ClusterRoles and bindings with new definitions.

## Edge cases and pitfalls
- `kubectl auth can-i --list` shows the full effective permission set; use it when the denial message is vague.
- Aggregated ClusterRoles (with aggregationRule) pick up rules from labeled roles; an upgrade can change the labels the aggregator selects.
- Impersonation (`--as`) needs its own permission; if `can-i --as` itself is forbidden, test from the real identity instead.
- Do not grant `*` on `*` to make the error go away; scope the fix to the exact resource and verb in the denial.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_L8HtIu-i_M7YRdqlzWI8Rw
