how to run a fleet of subagents without collisions
Explains how to run a fleet of subagents without collisions: partitioning work, claiming items atomically, per-worker directories, and idempotent retries. Use it when parallel agents share a queue, a filesystem, or an API. Triggered by questions about subagent coordination, parallel worker safety, or avoiding duplicate work. Not for single-agent workflows or for distributed systems theory.
TL;DR
Give every worker its own slice of the work, make claiming an item atomic (locks or atomic renames, never read-then-write), and design every task so a retry is harmless. Collisions come from shared mutable state; the fix is less shared state, claimed atomically, with idempotent work. A fleet that can be killed and restarted without corruption is a fleet you can trust.
how to run a fleet of subagents without collisionsUse this when
- Multiple subagents pull from one queue or work list
- Workers write to shared files or directories
- You have seen duplicate work, clobbered files, or lost updates
- You are designing a parallel pipeline from scratch
- A worker died mid-task and you need safe recovery
Not for this skill when
- Only one agent does the work (no coordination problem exists)
- You need formal distributed consensus (this is practical fleet ops, not theory)
- The shared resource is a database with real transactions (use the database's locking)
- Workers never share state at all (then there is nothing to collide on)
Steps
- Partition the work so workers rarely need the same thing. Assign id ranges, shard by hash, or hand out explicit chunks at spawn time. Partitioning is cheaper than locking; a fleet that never contends never collides.
worker A: items 1-50, worker B: items 51-100Expected output: disjoint assignments. Success check: two workers' logs show zero overlapping item ids.
- Make claiming atomic where partitioning is not enough. Use file locks, atomic renames, or a claim field updated in one locked operation. Read-then-write claiming (check status, then set status) races; two workers will read "pending" simultaneously and both will work the item.
flock queue.lock -c 'python3 claim_next.py --worker A'Expected output: each item claimed by exactly one worker. Success check: the results log shows no item processed twice.
- Give each worker its own scratch space. Per-worker temp directories, per-worker log files, per-worker output prefixes. Shared scratch directories are where "mystery" overwrites come from; separation makes every write attributable.
/tmp/fleet/worker-A/, /tmp/fleet/worker-B/Expected output: isolated working areas. Success check: deleting one worker's directory cannot affect another's.
- Make every unit of work idempotent. Claiming the same item twice, running the same step twice, or restarting mid-item must all be safe: skip-if-done checks, deterministic output names, and append-only logs. Idempotency is what lets you kill and restart workers without fear.
[ ] re-running a completed item is a no-op
[ ] partial outputs are overwritten deterministically, not appended twiceExpected output: safe retries everywhere. Success check: kill a worker mid-run, restart it, and verify zero corruption and zero duplicates.
- Log claims, completions, and heartbeats centrally. One append-only log per fleet with worker id, item id, and timestamps turns "what happened" from archaeology into reading. Monitor the log for stalls, not just failures.
[timestamp] worker-A claimed item 42
[timestamp] worker-A completed item 42Expected output: a central log. Success check: you can reconstruct any item's history from the log alone.
Variant phrasings
Parallel subagents stepping on each other
Problem phrasing. Steps 1 and 2: partition first, atomic claims where you cannot partition.
File lock for agent workers
Mechanism phrasing. Step 2's flock pattern is the minimal viable lock; use it around the smallest critical section that stays correct.
How to shard work across agents
Sharding phrasing. Step 1 expanded: by id range, by hash, or by explicit assignment, chosen for your workload's shape.
Safe retries for agent pipelines
Retry phrasing. Step 4: idempotent units plus skip-if-done checks make retries safe by construction.
Why it happens
Subagents are processes, and processes sharing state race. The classic failure is read-then-write: both workers see "pending", both act, and the second one's write silently wins or corrupts. Partitioning removes the shared state; atomic claims serialize the unavoidable sharing; idempotency makes the remaining races harmless. Every collision story is one of these three missing.
Edge cases / pitfalls
- Locks around long operations serialize your fleet. Keep the locked section tiny (the claim), never the work itself.
- A crashed worker can hold a claim forever. Claims need timeouts or heartbeats so a dead worker's items return to the pool.
- Filesystem atomicity varies. Rename is atomic on POSIX; check your platform before relying on it, and prefer real locks for the claim step.
- Do not let workers share API credentials with rate limits. Either partition the quota or serialize the calls; shared keys with shared limits are a collision in disguise.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_3x8hbb5GzbX1vmGWg6hQng
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.