agent's shutdown schedule stopped the canary instances - it matched the name pattern "test-" that prod canaries also...
Fixes a shutdown schedule that stops canary instances because it matches the 'test-' name prefix that production canaries also use. Use when scheduled stop/start automation hits the wrong instances. Key trigger: the agent selects targets by name pattern instead of an explicit opt-in tag.
TL;DR: Never select shutdown targets by name substring. Switch the schedule to an exact tag such as shutdown-schedule=true, and give canaries an explicit exclusion tag the agent always honors. Names describe things; tags declare intent. The agent confused the two.
Stopping instances matching name pattern "test-*": i-0aaa111 (prod-canary-1), i-0bbb222 (prod-canary-2), i-0ccc333 (test-runner-9)- Get the canaries back: run
aws ec2 start-instances --instance-ids i-0aaa111 i-0bbb222. Expected: instances return to running, health checks pass, the deployment pipeline unblocks. - Replace the name filter with a tag filter:
aws ec2 describe-instances --filters Name=tag:shutdown-schedule,Values=true Name=instance-state-name,Values=running. Expected: only instances explicitly opted in are returned. - Add an exclusion tag the agent checks first: if an instance carries do-not-stop at all, skip it no matter what else matches. Expected: canaries tagged once are protected from every schedule forever.
- Put a sanity gate on the schedule: the dry run must print the full target list, and it must refuse to run if the count exceeds your expected maximum or if any instance carries a prod tag. Expected: a bad filter fails loudly before stopping anything.
- Audit every other schedule and script for name-pattern matching on start/stop/terminate. Expected: zero remaining substring filters on names for destructive actions.
Use this when
- A scheduled stop/start killed or stopped the wrong instances
- Your agent picks instances by name pattern or wildcard
- Canary, staging, or shadow-prod instances share a naming prefix with real test boxes
- You are setting up automated shutdown schedules and want safe targeting
Not for this skill when
- The wrong instances were stopped because of a wrong REGION, not a wrong filter (different problem)
- The schedule timing itself is wrong (bad cron expression) rather than the targeting
- The instances failed to stop due to permissions (that is an IAM problem, not a targeting problem)
Variant phrasings
- shutdown script stopped production instances
- agent stopped instances matching test name
- canary instances shut down by schedule
- scheduled stop hit the wrong EC2 instances
Why it happens
Names are human-readable and easy to glob, so agents reach for them first. But naming conventions drift: prod canaries get 'test-' names precisely because they test production. A name describes what something looks like; a tag declares what should happen to it. The agent treated a description as an instruction.
Edge cases
- Instances with no tags at all: decide the default explicitly (opt-in is safer for destructive schedules) and document it
- Tags must be applied at provisioning time: any instance launched without the tag slips through, so put tagging in the launch template or Terraform
- Stopping is not free of side effects: instance-store data is lost and public IPs change on stop/start unless you use Elastic IPs
- Some instances cannot be stopped (instance-store backed): the agent should catch that error and report it, not crash the whole schedule
- ASG-managed instances will just get replaced when stopped: exclude them from stop schedules or suspend the ASG first
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_s9jS6BCEzekbXmAX9lwpuw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.