TL;DR: Restore the GPU capacity first, then give the agent the training schedule so it stops treating between-run idle time as waste. GPU nodes are the most expensive idle capacity in the account, so agents target them first - but a node at zero utilization for 11 weeks can be the most important machine in the company for one weekend. Only a schedule, not utilization, can tell the difference.

```text
cost agent killed the "idle" GPU nodes  -  it didn't know training jobs only run on the last weekend of the quarter
```

1. Check whether the killed nodes can come back. Look at whether they were stopped or terminated, and whether the training job definitions and data still exist. Expected: you know within minutes whether this is a restart or a rebuild.
2. Restore the capacity first and investigate second. Relaunch the GPU nodes from the launch template, then confirm the next quarterly training run has its targets. Expected: the upcoming training window is covered before you touch the agent's logic.
3. Give the agent the training schedule. Jobs that run on a cadence (end of quarter, monthly retrain) go into a schedule file the agent reads before any GPU action. Expected: the agent lists the next scheduled job when it evaluates GPU nodes.
4. Change the idle rule for GPUs: a GPU node is only a termination candidate if no job has run on it in the last two full schedule cycles AND no job is scheduled in the next cycle. Expected: quarterly training nodes never appear on the kill list.

## Use this when
- A cost agent stopped or terminated GPU nodes that run periodic training jobs
- GPU utilization reads zero for weeks between training runs
- You run ML training on a cadence (quarterly, monthly, weekly) rather than continuously

## Not for this skill when
- The GPUs are genuinely abandoned: no jobs scheduled, no recent runs, team confirms
- The nodes were only stopped (not terminated) and jobs can resume - restart them and add the schedule guard
- Training runs on spot or ephemeral clusters that are supposed to be torn down - check the cluster lifecycle policy instead

## Variant phrasings
- cost optimizer terminated gpu instances used for quarterly training
- agent flagged idle gpus that run periodic ml jobs
- how to exclude scheduled training nodes from idle cleanup
- finops agent killed p4 instances between training runs

## Why it happens
GPU nodes are the most expensive idle capacity in most accounts, so agents target them first. But idle now and unused are different things for batch ML: a node can sit at zero utilization for 11 weeks and then be the most important machine in the company for one weekend. Utilization-based rules cannot see the future; only a schedule can.

## Edge cases
- Ad-hoc training breaks the cadence rule. Require data scientists to register ad-hoc runs, or keep a small warm pool outside the agent's reach.
- Terminated nodes may have held local checkpoints. Check whether checkpoints were on EBS or instance storage before assuming recovery is clean.
- Multi-team GPU pools need per-team schedules. One team's quarterly job is another team's idle node.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_svL96Pl4B8lncaUn90h8kw
