agent's dry-run flag was inverted - it ran the deletion plan it was supposed to preview
Diagnoses a cost-agent cleanup script whose dry-run flag logic was inverted, so it executed the deletion plan it was meant to preview. Use when a FinOps automation deleted, stopped, or resized resources during what should have been a read-only run. Key trigger: the run log claims preview mode while CloudTrail shows mutating API calls in the same window.
TL;DR: Freeze the automation first by disabling its schedule trigger, then build the exact blast list from CloudTrail Delete, Terminate, and Stop calls inside the run window. The inversion is almost always a double negative in the flag plumbing, so rewrite the guard as an explicit opt-in to live mode with dry run as the default. Re-run the preview and confirm zero mutating API calls before anything goes live again.
Verbatim signal from the incident:
[2026-10-07 02:14:03] DRY-RUN: would terminate i-0a1b2c3d4e5f60718
[2026-10-07 02:14:04] DRY-RUN: would terminate i-0a1b2c3d4e5f60719The log said preview. CloudTrail showed real TerminateInstances calls one second later.
- Disable every trigger that can fire the script, EventBridge rules and cron entries alike, so it cannot re-fire while you investigate.
Command: aws events disable-rule --name [RULE-NAME] --region [REGION] Expected: the rule state reads Disabled on describe-rule.
- Pull the real blast list from CloudTrail for the run window, one region at a time since lookup is regional.
Command: aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventName,AttributeValue=TerminateInstances --start-time 2026-10-07T02:00:00Z --end-time 2026-10-07T03:00:00Z --region [REGION] Expected: a concrete list of affected resource IDs with timestamps, which you diff against the preview list in the logs.
- Repeat step 2 for StopInstances, DeregisterImage, DeleteSnapshot, and DeleteVolume in every region the script touches.
Expected: no region left unexamined, so the blast list is complete.
- Find the inversion in the flag parsing. The classic shape is a default-live variable combined with a flag that negates it a second time, or an argparse recipe copied wrong. Rewrite it so live mode requires an explicit flag and everything else previews:
parser.add_argument("--live", action="store_true", help="actually perform the changes")
args = parser.parse_args()
live_mode = args.liveExpected: running with no flags can only ever preview.
- Run the fixed script with no flags and watch CloudTrail for the new run window.
Expected: the plan prints, and lookup-events returns zero mutating calls.
- Restore what was destroyed from snapshots, AMIs, or backups, then re-enable the schedule only after a witnessed clean preview.
Expected: restored resources verified running before the automation is re-armed.
Use this when
- Resources were deleted, stopped, or resized during a run that was supposed to preview only
- Logs claim dry-run or preview mode but CloudTrail shows TerminateInstances, StopInstances, or Delete calls in the same window
- The flag uses a double negative, a store_false recipe, or an inherited boolean you do not trust
- You inherited a cleanup script and need to prove the preview path is safe before scheduling it
Not for this skill when
- The error is AWS's own DryRunOperation response to a DryRun parameter, which is the API confirming permissions, not a script bug
- Resources were deleted by a different automation, a human in the console, or an expired lifecycle policy
- The script failed closed and changed nothing, which is a permissions issue instead
Variant phrasings
- dry run deleted resources anyway
- script ran in live mode despite the dry run flag
- --dry-run flag was ignored and instances got terminated
- preview mode deleted EC2 instances
Why it happens
Boolean plumbing with two negatives reads correctly and runs wrong. A variable named live defaults to true, a --dry-run flag sets a second variable, and the guard combines them in a way that cancels out. Copy-pasted argparse recipes make it worse: action="store_false" on a flag already named as a negative flips the meaning silently, and nothing in the log output reveals it because the log line prints before the branch.
Edge cases
- CloudTrail lookup-events covers only the last 90 days and is per region, so a cross-region script needs one query per region
- The script may have two triggers, such as an EventBridge rule plus a cron entry on an instance, and disabling one still leaves the other live
- Some deletions are irreversible: terminated instance store volumes and deleted snapshots with no AMI are gone, so restore planning starts from backups, not from wishful thinking
- A second inverted flag can hide in a helper module the main script imports, so grep the whole repo for dry_run, preview, and live before declaring it fixed
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_V6Til0eDBGVClXdns-S-KQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.