profiling agent couldn't reproduce the slowdown -- the issue only happens under production traffic patterns the agent...
Helps a profiling agent reproduce production-only slowdowns it cannot trigger in test. Use when an agent's isolated benchmark looks clean but production latency is bad. Key trigger: the slowdown depends on traffic patterns, shared resources, or data shapes the test environment does not replicate.
Profiling agent cannot reproduce the production-only slowdown
TL;DR: Stop trying to reproduce it in isolation and instead capture the production conditions: traffic shape, concurrency, data volume, and noisy neighbors. Reproduce the environment, not just the endpoint. Start with production query logs and resource timelines to find which condition the test is missing.
profiling agent couldn't reproduce the slowdown -- the issue only happens under production traffic patterns the agent can't simulateSteps
- Characterize the production incident first - when, how often, how long:
-- from your slow query log or APM: distribution of the slow endpoint's latency
SELECT date_trunc('hour', ts) AS h, percentile_cont(0.99) WITHIN GROUP (ORDER BY ms) AS p99
FROM request_log WHERE endpoint = '/slow-one' AND ts between now() - interval '7 days' and now()
GROUP BY 1 ORDER BY 1;Expected: the p99 spikes cluster at specific hours or days - that pattern is your first reproduction clue.
- Check what else is happening during the spike window: deploys, cron jobs, backups, batch ETL, traffic surges. Correlate the spike hours with your deploy log and cron schedule.
Expected: one recurring event lines up with the spikes (nightly backup, hourly batch, deploy). That is the missing condition.
- Compare the data the test uses against production: row counts, table sizes, index bloat, statistics age:
SELECT relname, n_live_tup, last_analyze, last_autoanalyze
FROM pg_stat_user_tables WHERE relname = 'your_table';Expected: production has 100x the rows, stale stats, or heavy bloat the test database lacks.
- Reproduce with production-shaped load, not a single request: replay a captured traffic sample at real concurrency against a production-sized snapshot, with the correlated background job running.
Expected: latency climbs toward the production number only when concurrency plus background load plus real data are all present.
- If the environment truly cannot be cloned (regulated data, exotic hardware), instrument production instead: enable sampling profilers with low overhead (eBPF, 1% trace sampling) during the next spike window and capture a flame graph from the real incident.
Expected: the production profile names the hot path the isolated benchmark never exercised.
- Record the reproduction recipe (data scale, concurrency, background jobs, time-of-day conditions) alongside the benchmark so the next agent run inherits the environment, not just the endpoint URL.
Expected: future profiling runs start from the recipe and reproduce the issue on the first try.
Use this when
- An agent profiled an endpoint in isolation and found nothing, but production is slow.
- The slowdown correlates with time of day, deploys, or background jobs.
- Test data is smaller, fresher, or cleaner than production data.
- You need a systematic way to find the missing reproduction condition.
Not for this skill when
- The slowdown reproduces in isolation too - then it is a plain code or query problem, profile it directly.
- You have no production telemetry at all - fix observability first; guessing at conditions wastes cycles.
- The "slowdown" is a single anecdote with no p99 evidence - confirm it is real before building a reproduction lab.
Variant phrasings
- "can't reproduce production latency in staging"
- "endpoint slow in prod but fast in test"
- "agent benchmark clean but production p99 high"
- "how to reproduce production traffic patterns for profiling"
Why it happens
Latency is a property of the system under load, not of the code in isolation. Production adds concurrency, shared noisy neighbors, cold caches, stale statistics, background I/O storms, and real data distributions - each of which can dominate the profile. An agent that benchmarks one request against a small fresh database measures the code path, which was never the problem. The slowdown lives in the conditions around the code.
Edge cases
- Replaying production traffic can leak PII into test - scrub or synthesize the replay data.
- Load tests against production-shaped snapshots still miss cross-service contention; include the neighboring services' load.
- Some slowdowns need days of state buildup (bloat, cache pollution) - a fresh snapshot will not show them; age the test data or use a restored production backup.
- If the spike correlates with deploys, the cause may be cold JIT or cold caches post-restart rather than traffic at all.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_O7fGi0nb1MBMPUsBtkQJPw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.