## TL;DR

The agent did the right analysis against the wrong region because `us-east-1` was baked into the script while production lives in `eu-west-1`. Make the region an explicit, validated input on every cloud API call, verify the change set against the target region's inventory before acting, and add a guardrail that refuses to modify resources outside the declared target region.

## The query

```text
cost agent's rightsizing script resized instances in the wrong region - it hardcoded us-east-1 while prod runs in eu-west-1
```

## Use this when

- Cost automation modified resources in the wrong region.
- A region value is hardcoded anywhere in the automation.
- You need guardrails before the next rightsizing run.

## Not for

- Rightsizing math errors within the correct region (wrong instance type, bad utilization window).
- Actions blocked by IAM permissions (that is an authz problem, not a region problem).
- Multi-region fleets where acting in several regions is intended.

## Steps

### Step 1: Assess and revert what changed in the wrong region

```bash
aws ec2 describe-instances --region us-east-1
```

Expected output: the us-east-1 inventory. Identify the instances the script touched (compare against the pre-run inventory or CloudTrail ModifyInstanceAttribute events). Revert each to its previous instance type.

### Step 2: Verify the real production region is untouched

```bash
aws ec2 describe-instances --region eu-west-1 --query "Reservations[].Instances[].[InstanceId,InstanceType,State.Name]"
```

Expected output: the eu-west-1 inventory, unchanged. Confirm counts and types match the pre-run inventory. If anything did change there, it reverts the same way as step 1.

### Step 3: Parameterize the region, never hardcode it

```python
# region comes from config, validated at startup, never a literal
TARGET_REGION = config.require("target_region")  # e.g. eu-west-1
 ec2 = boto3.client("ec2", region_name=TARGET_REGION)
```

Expected output: every cloud client in the script is constructed with the configured region. A grep for `us-east-1` (or any region literal) in the automation now returns nothing.

### Step 4: Add a pre-flight check that prints the blast radius

```bash
echo "target region: eu-west-1, account: [account], instances in scope: [n]"
```

Expected output: before any mutating call, the script prints the region, account, and the exact count of resources in scope, and waits for confirmation on production. A wrong region is now visible before anything changes, not after.

### Step 5: Add a guardrail that aborts on out-of-region resources

```python
# refuse to touch anything outside the declared target region
for r in change_set:
    assert r["region"] == TARGET_REGION, "out-of-region resource blocked"
```

Expected output: the run aborts loudly if any planned change targets a resource outside the declared region. The guardrail turns a silent wrong-region run into a loud no-op.

### Step 6: Re-run the analysis scoped to the correct region

```bash
python rightsizing.py --region eu-west-1 --dry-run
```

Expected output: the rightsizing recommendations are computed against eu-west-1 utilization data, and the dry run shows exactly what would change. Apply only after reviewing the dry-run report.

## Variant phrasings

### agent's shutdown schedule stopped the canary instances
Same wrong-scope problem in the name dimension. The pre-flight check in step 4 lists instances by name pattern before acting.

### cost agent stopped the wrong RDS instances
Substring tag matching hit prod. Step 5's guardrail generalizes: assert every target matches the intended scope, whether region, tag, or name.

### rightsizing applied to dev instead of prod
Inverted scope, same fix. Explicit target plus pre-flight print plus abort-on-mismatch.

## Why it happens

SDKs default to a region, and a hardcoded default survives every copy-paste of the script across environments. The agent read the config it was given (us-east-1), did careful analysis, and acted confidently. Nothing in the loop ever asked "is this the region I meant." Region is ambient context to a human and an invisible default to a script, which is why it must be explicit, printed, and guarded rather than assumed.

## Edge cases

- Some resources are global (IAM, CloudFront, Route 53). The guardrail in step 5 must allowlist genuinely global resources instead of blocking them.
- The AWS CLI and SDKs each have their own default-region chains (env var, config file, IMDS). Pin the region explicitly in code. Do not rely on environment inheritance.
- A dry run that prints the region is only useful if a human reads it. For fully autonomous runs, the assert in step 5 is the real protection. The print is for the audit log.
- Reverting an instance type does not revert the billing for the hours it ran at the wrong size. Note the cost impact in the incident record.
- If the script ran against multiple regions in a loop, check every region in the loop, not just the two named here.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_gd8IVoldDRJZsyrauCiOww
