# Memory spike only on the first request after deploy - cold JIT

**TL;DR:** Profile a cold start, not a warm server. The first requests after a deploy pay JIT compilation, class loading, and cache priming all at once, which spikes memory before it settles. Capture a heap profile during the deploy window itself, or warm the process before it takes traffic.

```text
agent couldn't reproduce the memory spike because it only happens on the first request after a deploy when the JIT is cold
```

## Steps

1. Confirm the spike lines up with deploys: overlay container memory (or JVM heap) against deploy timestamps over the last few releases.
   Expected: a sharp memory peak within minutes of each deploy, decaying to baseline - the signature of cold-start allocation.

2. Capture what the cold path actually allocates. Take a heap dump or allocation profile during the first minute after a deploy (JFR allocation profiling, async-profiler alloc mode, or py-spy for Python equivalents):
   ```sh
   # example: start allocation profiling at boot, keep it short
   async-profiler -e alloc -d 60 -f /tmp/cold-alloc.html PID
   ```
   Expected: the profile is dominated by class loading, JIT-compiled code caches, and first-touch initialization - not your request handlers.

3. Compare against a warm profile of the same endpoint taken an hour later.
   Expected: the warm profile shows none of the cold-start allocation; the delta is the deploy spike, quantified.

4. Reduce the spike at the source: pre-warm the process before it receives traffic (health-check-gated warmup requests, readiness probes that wait for JIT), raise initial heap/code-cache sizes so the first expansion is not a cliff, and defer heavy lazy singletons to build time where possible.
   Expected: post-deploy peak shrinks and the process reaches steady state faster.

5. If the spike still threatens the container limit, add headroom: size memory limits for the cold peak, not the warm baseline, or roll deploys more gradually so fewer cold instances spike at once.
   Expected: no more OOM kills during rolling deploys.

6. Change the agent's profiling recipe: any memory investigation must include one cold-start capture (fresh process, first requests) alongside the warm profile.
   Expected: cold-start allocation stops being invisible to the agent.

## Use this when

- Memory spikes or OOM kills happen right after deploys and never in steady state.
- The agent profiled a long-running warm process and found nothing.
- You run JVM, .NET, or any runtime with JIT and lazy initialization.
- You need to decide between warming, tuning, or adding headroom.

## Not for this skill when

- Memory grows steadily over hours or days - that is a leak, not a cold start; profile the warm process over time.
- The spike happens on every request, not just post-deploy - look at per-request allocation instead.
- You run an interpreted runtime with no JIT - cold-start costs there are import and cache-priming, a smaller but similar investigation.

## Variant phrasings

- "memory spike after deploy then settles"
- "jvm heap spikes on first request after restart"
- "can't reproduce memory spike, only happens in production after release"
- "cold start memory allocation profiling"

## Why it happens

A fresh runtime process is expensive to bring up: the JIT compiles hot methods on first use, classes load and verify, static initializers run, and every cache starts empty so the first requests allocate what later requests reuse. Benchmarks and agents profile warm processes because that is what is running when they attach - the cold window lasts minutes and is gone by the time anyone looks. Production, however, creates a fresh cold process on every deploy.

## Edge cases

- Rolling deploys multiply the spike: N new pods each peak at once - stagger or pre-warm.
- Readiness probes that return too early let traffic hit a still-cold process, turning a memory spike into a latency spike too.
- Native-image / AOT builds remove JIT warmup but keep cache-priming allocation - the spike shrinks, it does not vanish.
- Allocation profilers add overhead; keep the cold capture short or it distorts the very startup you are measuring.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_OH77r9bi7Agl7QTFxrDEDg
