SRE team topology for agent-assisted ops
Defines SRE team topology for agent-assisted ops: a platform team owns the agent harness, identities, guardrails, and evals, while service teams own runbooks and escalation paths, with humans always on call and agents as scoped responders. Use when introducing agents into an existing SRE org or clarifying ownership. Not for tiny teams or read-only assistants.
TL;DR
Agent-assisted ops works best with a platform team owning the agent harness (identities, guardrails, evals) and service teams owning the domain knowledge (runbooks, escalation paths). On-call stays human; the agent is a responder with scoped powers, not a replacement. Write down who owns what before the first incident, not during it.
Error / query
SRE team topology for agent-assisted opsUse this skill when
- you are introducing agents into an existing SRE org
- agent responsibilities overlap confusingly with human on-call
- multiple teams want agents with different scopes
- leadership asks for a RACI covering agent operations
Not for this skill when
- the team is three people (keep it informal)
- agents are read-only assistants (lighter structure works)
- you are reorganizing the whole eng org (bigger question)
Steps
Step 1: Stand up a platform team owning the agent harness
mkdir -p agent-platform/{identities,policies,evals,runbooks}
ls agent-platform/Expected: one home for everything agent-shaped. The platform team owns ServiceAccounts and IAM roles, the guardrail policies, the staging eval suite, and the audit pipeline. No service team mints its own agent identity; that is how sprawl starts.
Step 2: Make service teams own their agent's domain knowledge
ls runbooks/payments/
cat runbooks/payments/CODEOWNERS || echo "no owners file yet"Expected: runbooks and escalation paths owned by the team that owns the service. The platform team provides the harness; the service team provides the judgment encoded as runbooks, approval rules, and "never do this" lists. An agent without team-owned knowledge is just a faster way to be wrong.
Step 3: Define the on-call relationship explicitly
cat /opt/sre-docs/agent-oncall-contract.mdExpected: a short contract stating the agent pages no one, a human is always the incident commander, the agent acts only within its guardrails, and any agent action can be frozen by the on-call in one command. Write this before the first incident; negotiating it mid-outage goes badly.
Step 4: Give each team a scoped agent identity, not one shared bot
kubectl get serviceaccounts -n payments -l app.kubernetes.io/managed-by=agent-platform
kubectl get serviceaccounts -n search -l app.kubernetes.io/managed-by=agent-platformExpected: per-team ServiceAccounts with the platform label. Blast radius follows team boundaries; when the payments agent misbehaves, you freeze payments, not the whole fleet.
Step 5: Review the topology quarterly with incident data
grep -h "agent" /opt/postmortems/2026-q3/*.md | grep -i -c "guardrail\|approval\|handoff"
ls /opt/postmortems/2026-q3/ | wc -lExpected: a count of agent-related postmortem mentions against total incidents. If every incident mentions the handoff being confusing, the topology needs work; the org chart should evolve with the evidence.
Variant phrasings
"who owns the ai agent in sre org"
Platform team owns the harness and safety properties; service teams own the knowledge and the approval policies for their domain. Both own the postmortems for their layer.
"should agents be on the on-call rotation"
No. Agents do not carry the pager; they assist the human who does. Putting an agent in the rotation confuses accountability and pages nobody useful.
"scaling agent ops across many teams"
The platform team's job is making the safe path the easy path: one command to get a scoped agent, templates for runbooks and guardrails, evals as a service. Teams adopt because it is easier than rolling their own.
Why it happens
Without explicit topology, agent ownership defaults to whoever set it up, which means one enthusiastic engineer becomes the single point of failure for every team's automation. Centralizing the harness while distributing the knowledge gives you consistent safety properties without bottlenecking every team on one group, and it makes the blast radius match the org chart people already understand.
Edge cases and pitfalls
- Do not let the platform team become a ticket queue; self-service with guardrails beats manual provisioning.
- Service teams will under-invest in runbooks; make runbook coverage part of their operational readiness review.
- The on-call contract needs re-signing when agent capabilities change; new powers mean new terms.
- Watch for shadow agents: teams will build their own if the platform path is slow, so keep the official path fast.
- Rotating engineers through the platform team spreads agent-ops literacy; a siloed platform team becomes a mystery box.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_u1gc5vo7MPdINVmdZrD85Q
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.