cost agent killed the "idle" GPU nodes - it didn't know training jobs only run on the last weekend of the quarter
Recovers from and prevents a cost agent shutting down GPU nodes that look idle but run quarterly training jobs. Use when GPU capacity was stopped or terminated between scheduled training runs. It restores capacity first, gives the agent the training schedule as an allowlist, and gates GPU termination on scheduler history rather than current utilization.
TL;DR: Restore the GPU capacity first, then give the agent the training schedule so it stops treating between-run idle time as waste. GPU nodes are the most expensive idle capacity in the account, so agents target them first - but a node at zero utilization for 11 weeks can be the most important machine in the company for one weekend. Only a schedule, not utilization, can tell the difference.
cost agent killed the "idle" GPU nodes - it didn't know training jobs only run on the last weekend of the quarter- Check whether the killed nodes can come back. Look at whether they were stopped or terminated, and whether the training job definitions and data still exist. Expected: you know within minutes whether this is a restart or a rebuild.
- Restore the capacity first and investigate second. Relaunch the GPU nodes from the launch template, then confirm the next quarterly training run has its targets. Expected: the upcoming training window is covered before you touch the agent's logic.
- Give the agent the training schedule. Jobs that run on a cadence (end of quarter, monthly retrain) go into a schedule file the agent reads before any GPU action. Expected: the agent lists the next scheduled job when it evaluates GPU nodes.
- Change the idle rule for GPUs: a GPU node is only a termination candidate if no job has run on it in the last two full schedule cycles AND no job is scheduled in the next cycle. Expected: quarterly training nodes never appear on the kill list.
Use this when
- A cost agent stopped or terminated GPU nodes that run periodic training jobs
- GPU utilization reads zero for weeks between training runs
- You run ML training on a cadence (quarterly, monthly, weekly) rather than continuously
Not for this skill when
- The GPUs are genuinely abandoned: no jobs scheduled, no recent runs, team confirms
- The nodes were only stopped (not terminated) and jobs can resume - restart them and add the schedule guard
- Training runs on spot or ephemeral clusters that are supposed to be torn down - check the cluster lifecycle policy instead
Variant phrasings
- cost optimizer terminated gpu instances used for quarterly training
- agent flagged idle gpus that run periodic ml jobs
- how to exclude scheduled training nodes from idle cleanup
- finops agent killed p4 instances between training runs
Why it happens
GPU nodes are the most expensive idle capacity in most accounts, so agents target them first. But idle now and unused are different things for batch ML: a node can sit at zero utilization for 11 weeks and then be the most important machine in the company for one weekend. Utilization-based rules cannot see the future; only a schedule can.
Edge cases
- Ad-hoc training breaks the cadence rule. Require data scientists to register ad-hoc runs, or keep a small warm pool outside the agent's reach.
- Terminated nodes may have held local checkpoints. Check whether checkpoints were on EBS or instance storage before assuming recovery is clean.
- Multi-team GPU pools need per-team schedules. One team's quarterly job is another team's idle node.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_svL96Pl4B8lncaUn90h8kw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.