# Latency only appears when other services saturate the same node

**TL;DR:** Profile the node, not just the endpoint. When four services share a box, CPU steal, memory pressure, and disk I/O from neighbors show up as your service's latency. Correlate the endpoint's p99 with node-level saturation metrics, then isolate with resource limits or move the noisy neighbor.

```text
agent profiled the endpoint in isolation but the latency only appears when 4 other services saturate the same node
```

## Steps

1. Confirm the correlation: overlay the endpoint's p99 latency with node CPU, memory pressure, and disk I/O over the incident window (from your APM plus node exporter, CloudWatch, or equivalent).
   Expected: p99 spikes line up with node saturation (CPU above 85%, iowait high, or memory near limit) rather than with your service's own traffic.

2. Identify the noisy neighbor: rank processes or containers on the node by CPU and I/O during the spike:
   ```sh
   # on the node during (or from a captured snapshot of) the incident:
   ps -eo pid,comm,%cpu,%mem --sort=-%cpu | head -15
   ```
   Expected: a sibling service (batch job, indexer, backup agent) tops the list while your service sits far below.

3. Check for the classic shared-resource signatures: CPU steal on virtualized hosts, OOM-killer-adjacent memory pressure, and iowait from a neighbor's bulk writes. Container CPU throttling shows as `container_cpu_cfs_throttled_seconds_total` climbing in your metrics.
   Expected: at least one saturation signal coincides with the latency spike.

4. Reproduce honestly: run your endpoint benchmark while simultaneously loading the node the way the neighbors do (a CPU burner plus a disk writer at the observed levels), not on an idle box.
   Expected: the isolated-fast endpoint degrades to near-production latency - you now have a reproduction.

5. Fix it with isolation, in order of cheapness: set CPU/memory requests and limits so neighbors cannot starve your service; pin the batch workloads to specific cores or nodes; move the heaviest neighbor to its own node or schedule it outside peak hours.
   Expected: p99 drops and decouples from the neighbor's schedule.

6. Add node-saturation metrics to the agent's profiling checklist so isolated-endpoint benchmarks always ship with a "shared node state" section.
   Expected: future profiles note contention instead of declaring the endpoint clean.

## Use this when

- An endpoint is fast in isolation but slow in production on shared infrastructure.
- Latency spikes correlate with other services' activity, not your own traffic.
- You run multiple services per node, VM, or Kubernetes worker.
- You need to prove contention before asking for isolation or new capacity.

## Not for this skill when

- The service has dedicated nodes and still slows down - look at its own code, queries, or downstream dependencies.
- Latency is constant rather than spiky - steady slowness is a code or query problem, not contention.
- You cannot get node-level metrics - without them this stays a hypothesis; fix observability first.

## Variant phrasings

- "service slow only when other services busy same node"
- "noisy neighbor latency kubernetes"
- "endpoint fast alone slow in production shared host"
- "how to prove cpu contention from sibling service"

## Why it happens

Profilers attribute time to code, but wall-clock latency includes time the process spends waiting: descheduled while a neighbor burns CPU, stalled on disk behind a neighbor's writes, or throttled by the container runtime. An isolated benchmark removes all of that and measures a fantasy. The agent did the profiling right and the scoping wrong - it profiled the service instead of the system the service runs in.

## Edge cases

- CPU limits in Kubernetes can throttle your own service before the node saturates - check throttling metrics before blaming neighbors.
- Memory pressure causes page-cache eviction that looks like slow disk; the fix may be more RAM, not less neighbor.
- Network contention on shared hosts shows up as latency with clean CPU - check retransmits and interface saturation too.
- Moving the neighbor helps until traffic grows into the new headroom; set alerts on node saturation so the next collision gets caught early.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_CBNU3tXTw3B8kuHUwhMwyw
