Check the agent logs and host metrics for networking, CPU, memory, and I/O pressure around the time heartbeats stopped. On the Elastic CI Stack for AWS with spot instances, abrupt termination also marks agents lost; use the log collector script to gather evidence. Also check the command step timeout: a job exceeding it can be killed ungracefully and surface as -1 depending on the cancel-signal-timeout.

Context: Official Buildkite agent lifecycle docs: a -1 exit code means Buildkite lost contact with the agent. Agents send heartbeat updates after registering; if none arrive for three consecutive minutes, the agent is marked lost and gets no more jobs. Common causes are networking issues and CPU, memory, or I/O constraints on the agent host.