# TL;DR
The API response is the truth; the console page is a rumor. Capture the stop call's response, confirm the instance actually entered the stopping state with describe-instances, and let any error fail loudly. An agent that swallows a permission error will report victory while the meter keeps running.

```text
agent verified the shutdown worked by checking the console  -  the stop call hit a permission error it swallowed and the instances kept running
```

## Steps

1. Reproduce the call and watch the response instead of the console:
```
aws ec2 stop-instances --instance-ids [instance-id]
```
Expected: either a StoppingInstances block with the state moving to stopping, or a ClientError naming the missing permission.

2. Confirm with a second read:
```
aws ec2 describe-instances --instance-ids [instance-id] --query 'Reservations[].Instances[].State.Name'
```
Expected: "stopping" or "stopped". Anything else means the action did not land.

3. Find the swallowed error in the agent's code: look for catch-all handlers around the stop call that log and continue, and for "verify" steps that reload a cached page instead of re-querying.
Expected: you find the exact line where the error died silently.

4. Change the pattern: no error is caught without re-raising or halting the run, and verification always re-reads state from the API.
Expected: the next permission failure pages someone instead of printing "shutdown complete".

## Use this when
- instances still run after automated shutdowns
- the agent's "verified" claims do not match reality
- errors vanish into logs nobody reads

## Not for this skill when
- the stop call genuinely succeeded (check state first)
- the instances are being restarted by an autoscaler (that is a different fight)
- the error is throttling rather than permissions (add backoff instead)

## Variant phrasings
- "stop-instances permission error ignored"
- "agent said instances stopped but they are running"
- "automation swallowed AWS error"

## Why it happens
SDK calls surface failures as exceptions or error fields, and hurried automation wraps them in catch-all handlers that log at debug and move on. The "verification" step then checks a stale source (a console page, a cached inventory) instead of live state, so the lie survives all the way to the report.

## Edge cases
- DryRun is your friend: run stop-instances with --dry-run first. A DryRunOperation success means permissions are fine.
- Some instances cannot stop (instance-store root volumes, or members of an autoscaling group). The API tells you, if you listen.
- In cross-account setups the error may be a trust-policy denial, which looks identical to a missing permission. Check both.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_PWVj-qbf25Q9j-SenjynOg
