agent profiled the endpoint in isolation but the latency only appears when 4 other services saturate the same node
Fixes misleading profiles taken on an isolated endpoint when the real latency comes from node-level contention with other services. Use when production is slow but single-service benchmarks look fine. Key trigger: latency appears only when sibling services saturate the shared node.
Latency only appears when other services saturate the same node
TL;DR: Profile the node, not just the endpoint. When four services share a box, CPU steal, memory pressure, and disk I/O from neighbors show up as your service's latency. Correlate the endpoint's p99 with node-level saturation metrics, then isolate with resource limits or move the noisy neighbor.
agent profiled the endpoint in isolation but the latency only appears when 4 other services saturate the same nodeSteps
- Confirm the correlation: overlay the endpoint's p99 latency with node CPU, memory pressure, and disk I/O over the incident window (from your APM plus node exporter, CloudWatch, or equivalent).
Expected: p99 spikes line up with node saturation (CPU above 85%, iowait high, or memory near limit) rather than with your service's own traffic.
- Identify the noisy neighbor: rank processes or containers on the node by CPU and I/O during the spike:
# on the node during (or from a captured snapshot of) the incident:
ps -eo pid,comm,%cpu,%mem --sort=-%cpu | head -15Expected: a sibling service (batch job, indexer, backup agent) tops the list while your service sits far below.
- Check for the classic shared-resource signatures: CPU steal on virtualized hosts, OOM-killer-adjacent memory pressure, and iowait from a neighbor's bulk writes. Container CPU throttling shows as
container_cpu_cfs_throttled_seconds_totalclimbing in your metrics.
Expected: at least one saturation signal coincides with the latency spike.
- Reproduce honestly: run your endpoint benchmark while simultaneously loading the node the way the neighbors do (a CPU burner plus a disk writer at the observed levels), not on an idle box.
Expected: the isolated-fast endpoint degrades to near-production latency - you now have a reproduction.
- Fix it with isolation, in order of cheapness: set CPU/memory requests and limits so neighbors cannot starve your service; pin the batch workloads to specific cores or nodes; move the heaviest neighbor to its own node or schedule it outside peak hours.
Expected: p99 drops and decouples from the neighbor's schedule.
- Add node-saturation metrics to the agent's profiling checklist so isolated-endpoint benchmarks always ship with a "shared node state" section.
Expected: future profiles note contention instead of declaring the endpoint clean.
Use this when
- An endpoint is fast in isolation but slow in production on shared infrastructure.
- Latency spikes correlate with other services' activity, not your own traffic.
- You run multiple services per node, VM, or Kubernetes worker.
- You need to prove contention before asking for isolation or new capacity.
Not for this skill when
- The service has dedicated nodes and still slows down - look at its own code, queries, or downstream dependencies.
- Latency is constant rather than spiky - steady slowness is a code or query problem, not contention.
- You cannot get node-level metrics - without them this stays a hypothesis; fix observability first.
Variant phrasings
- "service slow only when other services busy same node"
- "noisy neighbor latency kubernetes"
- "endpoint fast alone slow in production shared host"
- "how to prove cpu contention from sibling service"
Why it happens
Profilers attribute time to code, but wall-clock latency includes time the process spends waiting: descheduled while a neighbor burns CPU, stalled on disk behind a neighbor's writes, or throttled by the container runtime. An isolated benchmark removes all of that and measures a fantasy. The agent did the profiling right and the scoping wrong - it profiled the service instead of the system the service runs in.
Edge cases
- CPU limits in Kubernetes can throttle your own service before the node saturates - check throttling metrics before blaming neighbors.
- Memory pressure causes page-cache eviction that looks like slow disk; the fix may be more RAM, not less neighbor.
- Network contention on shared hosts shows up as latency with clean CPU - check retransmits and interface saturation too.
- Moving the neighbor helps until traffic grows into the new headroom; set alerts on node saturation so the next collision gets caught early.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_CBNU3tXTw3B8kuHUwhMwyw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.