cost agent deleted EBS snapshots it thought were orphaned - the "owner: none" tag was a silent tagging failure at...
Recovers from and prevents a cost agent deleting EBS snapshots that looked orphaned because a silent tagging failure left them tagged owner none. Use after a snapshot cleanup deleted snapshots that were actually in use, or before running automated snapshot deletion. Not for deleting genuinely orphaned volumes, AMI deregistration, or snapshots confirmed unattached.
TL;DR
A snapshot tagged "owner: none" is not proof of being orphaned. It can be proof that the tagging step failed at creation. Before deleting any snapshot, verify it is not referenced by an AMI, a volume, or a backup policy, and require a positive orphan signal (no references plus age) rather than trusting the tag alone. Restore what you can, then fix the tagging pipeline so the failure is loud next time.
The query
cost agent deleted EBS snapshots it thought were orphaned - the "owner: none" tag was a silent tagging failure at creationUse this when
- A snapshot cleanup deleted snapshots that turned out to be in use.
- The deletion logic trusted an "owner: none" or missing-owner tag as the orphan signal.
- You are about to automate snapshot deletion and want the guardrails first.
Not for
- Deleting EBS volumes (different resource, different checks).
- Deregistering AMIs (the AMI lifecycle has its own ordering rules).
- Snapshots you have positively confirmed are unreferenced.
Steps
Step 1: Assess what was deleted and what referenced it
aws ec2 describe-images --owners self --query "Images[].BlockDeviceMappings[].Ebs.SnapshotId"Expected output: the snapshot IDs still referenced by your AMIs. Compare against the deletion log from CloudTrail. Any deleted ID that appears here broke an AMI, and that is the restore priority list.
Step 2: Restore from copies or backups where they exist
aws ec2 describe-snapshots --filters "Name=tag:source-snapshot,Values=[deleted id]" --region [dr region]Expected output: any cross-region copies or backup-vault copies of the deleted snapshots. Restore from those first. If no copy exists and the AMI is broken, rebuild the AMI from a running instance launched before the deletion.
Step 3: Replace the tag-trust deletion rule with a reference check
aws ec2 describe-snapshots --owner-ids selfExpected output: all your snapshots. Filter to candidates older than the cutoff date. For each candidate, the new policy requires all three before deletion: not referenced by any AMI, not created by a backup service within its retention window, and older than the cutoff. The tag is informational only, never the delete signal.
Step 4: Make snapshot tagging fail loudly at creation
Update the tagging policy: create-snapshot must verify tags after write, retry twice, and alert on failure.
Expected output: the creation pipeline now reads back the tags after writing them. If the tag write fails, creation fails loudly instead of silently producing an "owner: none" snapshot that a future cleanup will misread.
Step 5: Backfill owner tags from CloudTrail
aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventName,AttributeValue=CreateSnapshot --max-results 50Expected output: the original CreateSnapshot events with the requesting identity. Use them to backfill the true owner tag on surviving snapshots, so the tag data becomes trustworthy again.
Step 6: Add a human-reviewed dry run before any bulk deletion
Update the finops runbook: snapshot cleanup requires a dry-run report and human approval.
Expected output: the automation produces a report of exactly what it would delete (with reference-check results) and waits for approval. Bulk deletion without a preview is how this incident happened.
Variant phrasings
cost agent deregistered the wrong AMIs
Same trust-the-tag failure one level up. Check AMI references (running instances, launch templates) before deregistration, not just age tags.
agent deleted old snapshots that were the base of an active AMI chain
Incremental snapshots look old and unused while the whole chain depends on them. The reference check in step 3 catches this. The tag never will.
snapshot cleanup deleted everything with no owner tag
The cleanup treated missing data as a delete signal. Missing data is never a delete signal. Step 3 inverts the default to keep.
Why it happens
Tag-based cleanup assumes tags are truthful, but tagging is a separate API call that can fail silently (throttling, permissions, eventual consistency). When it fails, the resource gets a null or default owner value, and the cleanup agent reads that null as "nobody owns this, safe to delete." The signal the agent trusted was actually the error, laundered through a tag. Every automated deletion needs a positive liveness check against the real references, because tags describe intent and references describe reality.
Edge cases
- Snapshots shared with your account from another account do not appear under
--owner-ids self. The reference check must include shared snapshots or the cleanup will miss them. - Backup services (AWS Backup, DLM) manage snapshot lifecycles themselves. Deleting their snapshots out from under them breaks retention policies. Exclude backup-managed snapshots from custom cleanup entirely.
- A snapshot that is genuinely orphaned today may become an AMI base tomorrow. The age cutoff in step 3 exists to absorb that lag.
- CloudTrail lookup only goes back 90 days. For older snapshots, the backfill in step 5 cannot recover the original owner. Treat those as unknown-owner and apply the strictest reference check.
- Restoring a deleted snapshot from a cross-region copy preserves the data but not the original snapshot ID. Update any AMI or automation that referenced the old ID.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstpeEa9ewTj3ZtlRpJmU5zQ