red-teaming your own agents: the basics
A defensive starter guide to adversarially testing your own agents: scoped permission, attack-surface mapping, prompt-injection tests, scope-escape attempts, and exfiltration drills in a sandbox, then fix and re-test. Use before granting broad permissions or after an agent security incident. Triggers: 'red team AI agents', 'test agent prompt injection', 'adversarial testing agents'. Not for: attacking anyone else's systems, or evading someone's controls.
red-teaming your own agents: the basics
TL;DR
Attack your own agents before someone else does: get permission, define scope, then try the tricks attackers use, prompt injection, scope escapes, and exfiltration drills, against your own systems in a sandbox. Write up what worked, fix it, and test again. You are finding your weaknesses on purpose.
red-teaming your own agents: the basicsUse this when
- An agent is about to get broad permissions or handle sensitive data
- You want to know if your prompt-injection defenses actually work
- Leadership asks how you test agent safety beyond hoping
- After any agent security incident, to find the sibling weaknesses
Not for this skill when
- You want to attack someone else's agents or systems (that needs their explicit permission, full stop)
- You are looking for evasion techniques to bypass someone's controls (not what this is)
- The agent is a toy with no real permissions (the findings would not mean anything)
Steps
1. Get written permission and define scope.
Write down which agents, which systems, and which techniques are in play, and get sign-off from whoever owns them. Red-teaming without permission is just attacking. Scope keeps it a drill instead of an incident.
Expected: a one-page scope doc with signatures, or at least written approval in chat, before anything starts.
2. List what the agent can touch.
Inventory the agent's tools, credentials, data sources, and network reach. This is your attack surface map. You cannot test what you have not listed.
Expected: a list of every tool and credential the agent holds, which becomes your test checklist.
3. Try prompt injection against your own agent.
Feed it the attacks in a sandbox: a web page with hidden instructions, a document with directives buried in the middle, a pasted blob of text that tells it to ignore its task. See which ones it follows and which ones your defenses catch.
agent run --sandbox --task "summarize https://example.com/redteam-test-page"Expected: the agent summarizes the page and ignores or flags the hidden instruction. If it obeys the hidden instruction, you found a gap.
4. Try to escape the agent's scope.
Ask it, as an attacker would through indirect means, to reach outside its box: read another repo, use a credential for a different system, touch data outside its task. Every escape that works is a scoping failure to fix.
Expected: a list of escape attempts marked blocked or succeeded, with the successful ones filed as findings.
5. Try the exfiltration paths.
See if you can get the agent to send data somewhere it should not: an unapproved host, a pasted secret in chat, a bulk export. Test the egress allowlist and the approval gates under adversarial pressure.
Expected: every exfiltration attempt is blocked or gated, with each block in the logs.
6. Write up findings like a defender.
For each success, record what you tried, what the agent did, and the blast radius if a real attacker had done it. Rank by impact, not by cleverness. A boring scope escape beats a fancy injection every time.
Expected: a findings list with impact ratings and owners for each fix.
7. Fix, then re-test.
Patch the controls: tighten scoping, fix the sanitization, add the missing alert. Then run the same attacks again and confirm they fail. A finding without a re-test is a rumor.
Expected: every high-impact finding re-tested and confirmed fixed, with evidence.
Variant: test my agent for prompt injection
That is step 3 expanded: build a small library of injection test cases (hidden page text, poisoned docs, instruction-override pastes) and run them against every agent before it gets new permissions.
Variant: adversarial testing for AI agents
Same discipline, broader scope: add tool-abuse tests and multi-turn manipulation, where the "attacker" steers the agent over several steps instead of one shot.
Variant: how to pentest an AI agent
Scope it like any pentest: permission, targets, techniques, timeline, report. The agent-specific parts are the injection tests and the tool-call review; the rest is standard practice.
Variant: AI red team basics for small teams
You do not need a dedicated team. One engineer, one afternoon, this checklist, a sandbox. The findings from a first pass are always worth more than the effort.
Why this happens
Defenses that have never been tested are assumptions wearing uniforms. Filters, allowlists, and scope boundaries look solid on paper and frequently have gaps you only find by pushing on them. Attackers will push; red-teaming means you push first, with a fix ready.
Edge cases and pitfalls
- Testing in production by accident: use sandbox sessions with sandbox credentials, always. A red-team test touching real customer data is an incident you caused.
- The agent passes every test: either your defenses are good or your tests are weak. Have someone else write new test cases; familiarity breeds blind spots.
- Findings that never get fixed: assign owners and dates in the write-up. A known unfixed finding is worse than an unknown one, because now it is negligence.
- Scope creep mid-test: "while I am here" testing outside the scope doc is how drills become incidents. New idea, new permission, then test.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst9njRLLXKBIC8NL6RIgAFA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.