## TL;DR
ImagePullBackOff from ECR with IRSA almost always means the pod's service account is not getting valid AWS credentials: the OIDC provider is misconfigured, the IAM role trust policy does not match the service account, or the ECR helper cannot refresh. Check the service account annotation, the role's trust policy, and the exact pull error on the pod, then fix the trust relationship. The image itself is rarely the problem.

## Error / query
```text
ImagePullBackOff with ECR: IRSA token refresh issues
```

## Use this skill when
- Pods show `ImagePullBackOff` pulling from an ECR repository
- The cluster uses IRSA (service account annotated with an IAM role)
- The failure started after a cluster upgrade, addon change, or role edit
- `Failed to pull image` mentions ECR auth or `no basic auth credentials`

## Not for this skill when
- Pulling from Docker Hub, GHCR, or a non-AWS registry (different auth)
- Using static AWS access keys in a pull secret (check the secret instead)
- The error is `manifest unknown` or `not found` (wrong image name/tag)

## Steps

### Step 1: Read the exact pull error
```bash
kubectl describe pod [pod-name] -n [namespace] | grep -A5 "Failed to pull image"
```
Expected: the real reason, e.g. `no basic auth credentials` (auth never obtained) or `401 Unauthorized` (credentials obtained but rejected). Auth-never-obtained points at IRSA; 401 points at the IAM policy.

### Step 2: Verify the service account annotation
```bash
kubectl get serviceaccount [sa-name] -n [namespace] -o yaml | grep -i "role-arn"
```
Expected: an annotation like `eks.amazonaws.com/role-arn: arn:aws:iam::[account]:role/[role]`. Missing or wrong annotation means the pod never assumes a role; fix the annotation on the service account the pod actually uses.

### Step 3: Check the IAM role trust policy
```bash
aws iam get-role --role-name [role] --query Role.AssumeRolePolicyDocument
```
Expected: a trust policy allowing `sts:AssumeRoleWithWebIdentity` for your cluster's OIDC provider with a `StringEquals` on the service account (`system:serviceaccount:[namespace]:[sa-name]`). A mismatch here (wrong namespace, wrong account name) is the single most common cause.

### Step 4: Confirm the cluster OIDC provider exists and matches
```bash
aws eks describe-cluster --name [cluster] --query "cluster.identity.oidc.issuer"
aws iam list-open-id-connect-providers
```
Expected: the issuer URL from EKS matches a provider in IAM. After cluster recreations the issuer changes and the old provider entry goes stale; delete the stale one and create the current one.

### Step 5: Verify the IAM policy allows ECR pulls
```bash
aws iam simulate-principal-policy --policy-source-arn arn:aws:iam::[account]:role/[role] --action-names ecr:BatchGetImage ecr:GetDownloadUrlForLayer ecr:BatchCheckLayerAvailability --resource-arns arn:aws:ecr:[region]:[account]:repository/[repo]
```
Expected: all three actions allowed. If any is denied, attach the ECR pull permissions (or the managed `AmazonEC2ContainerRegistryReadOnly` policy) to the role.

## Variant phrasings

### "eks imagepullbackoff ecr"
Same flow. Steps 2 and 3 fix the majority of cases.

### "irsa ecr pull fails"
Focus on the trust policy (step 3) and OIDC provider (step 4); IRSA failures are identity problems, not image problems.

## Why it happens
With IRSA, the kubelet gets ECR credentials through a credential chain that depends on the pod's projected service-account token being accepted by AWS STS. Any break in the chain (wrong annotation, trust policy not matching the exact service account, stale OIDC provider, missing ECR permissions) leaves the kubelet with no credentials, and the pull fails before the image is ever contacted.

## Edge cases and pitfalls
- Region matters: the ECR repository region must match where the credential helper looks; cross-region pulls need the helper configured for that region.
- Service account changes require pod recreation; editing the SA annotation does not fix already-running pods, only new ones.
- The EKS Pod Identity Agent is a different mechanism from IRSA; if the cluster migrated, check which one is actually configured.
- `ImagePullBackOff` with a 30+ minute backoff can hide a fix; delete the pod to retry immediately after fixing auth.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_vLwkgOLYQlOydszNGIdQTA
