how to let an agent run kubectl safely
Gives an AI agent safe kubectl access via a dedicated ServiceAccount with a least-privilege Role, starting read-only and adding narrow write verbs only where needed, verified with auth can-i and covered by audit logging. Use when an agent investigates or remediates Kubernetes issues. Not for human operator access, cloud-only API needs, or bypassing RBAC.
TL;DR
Give the agent its own ServiceAccount with a narrowly scoped Role, never your admin kubeconfig. Start read-only, add write verbs only for the exact resources and namespaces it needs, and verify the boundaries with auth can-i before it runs anything. An agent with cluster-admin is an incident waiting for a bad suggestion.
Error / query
how to let an agent run kubectl safelyUse this skill when
- you want an AI agent to investigate or remediate Kubernetes issues
- an agent needs to read pod logs and events in specific namespaces
- you are setting up an agent that applies pre-approved manifests
- auditors ask what a service account used by automation can do
Not for this skill when
- a human operator needs kubectl access (use your identity provider and short-lived certs)
- the agent only needs cloud APIs, not the cluster (scope IAM instead)
- you want the agent to bypass RBAC or approval flows (do not do this)
Steps
Step 1: Create a dedicated ServiceAccount for the agent
kubectl create serviceaccount sre-agent -n ops
kubectl get serviceaccount sre-agent -n opsExpected: the ServiceAccount exists in the ops namespace. A dedicated account means you can audit, rotate, and revoke the agent's access without touching anything else.
Step 2: Bind a least-privilege Role, starting read-only
kubectl create role agent-reader --verb=get,list,watch --resource=pods,deployments,services,events,configmaps -n ops
kubectl create rolebinding agent-reader-binding --role=agent-reader --serviceaccount=ops:sre-agent -n opsExpected: both resources created. The agent can now observe the ops namespace but cannot change anything, which is the right starting posture for investigation agents.
Step 3: Add narrowly scoped write access only where needed
kubectl create role agent-restarter --verb=patch --resource=deployments/scale -n ops
kubectl create rolebinding agent-restarter-binding --role=agent-restarter --serviceaccount=ops:sre-agent -n opsExpected: the agent can scale deployments in ops and nothing else. Prefer subresource verbs (like deployments/scale) over broad update verbs; restarting via scale is safer than letting the agent edit specs.
Step 4: Verify the boundaries from the agent's perspective
kubectl auth can-i --list --as=system:serviceaccount:ops:sre-agent -n ops
kubectl auth can-i create pods --as=system:serviceaccount:ops:sre-agent -n kube-systemExpected: the first command lists exactly the verbs you granted; the second returns "no". If anything unexpected shows "yes", fix the binding before the agent runs.
Step 5: Enable audit logging for the ServiceAccount
kubectl get events -n ops --sort-by=.lastTimestamp | tail -20Expected: recent events visible. On top of this, make sure your cluster audit policy logs requests from system:serviceaccount:ops:sre-agent at RequestResponse level so every agent action is traceable (see the audit skill for the full setup).
Variant phrasings
"agent needs kubectl in production"
Same pattern, but keep prod to a separate ServiceAccount with even tighter verbs, require dry-run or approval for writes, and alert on every write the agent performs.
"how do i stop an agent from deleting things"
Do not grant delete at all. If a workflow genuinely needs cleanup, put deletes behind a human approval gate and let the agent propose the exact command instead of running it.
"share kubeconfig with an ai coding agent"
Do not hand it yours. Generate a kubeconfig bound to the scoped ServiceAccount only, with a short TTL, and rotate it regularly.
Why it happens
Agents act on probabilistic reasoning, so they will eventually misread a situation and run a destructive command with full confidence. RBAC is the blast-radius control: the damage an agent can do is bounded by the verbs you granted, not by how careful its instructions are. Least privilege turns "the agent might do anything" into "the worst case is a bad scale event in one namespace".
Edge cases and pitfalls
- Aggregated ClusterRoles can silently widen a namespaced Role; check for aggregation labels before binding.
- impersonation headers let a human act as the ServiceAccount; restrict who can impersonate it.
- Some agents cache credentials; rotating the ServiceAccount token does not help if the agent kept a copy in its working files, so scope the kubeconfig TTL short.
- Wildcard resources ("*") in a Role defeat the whole exercise; enumerate resources explicitly.
- Audit logs are only useful if someone reads them; pair this with alerting on agent write actions.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_A4nXOh54V9o8GZDvfwYHPg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.