## TL;DR

An N+1 detector tuned for request latency will always scream at a bulk import, because imports legitimately issue thousands of similar queries. Tag background jobs in the query telemetry (job name, queue, batch size) and scope N+1 rules to request-scoped traces only. A duplicate pattern inside a single job run is an import to optimize, not an endpoint emergency.

## The query

```text
agent's N+1 alert fired on a nightly bulk import job -- it treated a background job like a hot endpoint
```

## Use this when

- A duplicate-query or N+1 alert fires on a batch job, ETL run, or seed script.
- The alert has no request context (no endpoint, no user-facing latency).
- The profiler agent pages or files tickets for background workloads.

## Not for

- Genuine N+1 loops in request handlers serving user traffic.
- An import that is actually too slow and needs batch optimization (that is real, just a different rule).
- Alerts with no query sample attached (that is a telemetry gap, not a false positive).

## Steps

### Step 1: Confirm the trace is a background job

```bash
grep -o '"job[^,]*' trace-[alert-id].json | head -5
grep -o '"endpoint[^,]*' trace-[alert-id].json | head -5
```

Expected output: the trace carries job or queue tags and no HTTP endpoint tag. That confirms the detector fired on a background workload, which the hot-endpoint rule was never designed for.

### Step 2: Exclude job-scoped traces from the hot-endpoint N+1 rule

```yaml
n1_rule:
  scope: request_traces_only
  exclude_tags: ["job", "queue", "cron", "worker"]
```

Expected output: the N+1 rule stops evaluating traces that carry job tags. The nightly import no longer pages anyone. Request traces are unaffected.

### Step 3: Create a separate bulk-job rule with sane thresholds

```yaml
bulk_job_rule:
  scope: job_traces_only
  duplicate_threshold: 10000
  alert_on: job_duration_regression
```

Expected output: background jobs get their own rule that alerts on what actually matters for jobs (duration regression, throughput collapse), not on per-query duplication, which is normal for bulk work.

### Step 4: Add job identity to every query log line

```sql
-- tag queries at the connection level so every sample carries context
SET application_name = '[job name]:[batch id]';
```

Expected output: every query in the log now carries the job name and batch ID. The next alert arrives with context attached, and the agent can route it to the right rule without guessing.

### Step 5: Re-tune the detector to count per request, not per process

Record in the detector tuning notes: count duplicates per trace, not per process lifetime.

Expected output: the tuning note is recorded. A detector that counts duplicates across a whole long-lived worker process will always false-positive on jobs. Per-trace counting keeps the semantics tied to a single unit of work.

### Step 6: Verify the import is not genuinely pathological

```bash
grep -c "INSERT" import-query-log.sql
```

Expected output: a sanity count of what the import actually did. If the import issues one INSERT per row for a million rows, that is worth optimizing (bulk COPY, batch inserts) as a batch problem, on the job rule's terms, not as an N+1 emergency.

## Variant phrasings

### agent flagged an N+1 from the staging seed-data generator
Same workload-context miss. Seed scripts are jobs. Step 2's exclusion covers them.

### profiler reported 1200 queries per request but the trace included its own instrumentation
Same false-positive family: the detector counted things that are not app queries. Step 4's tagging plus an instrumentation exclusion fixes it.

### duplicate-query detector tripped on prepared statements with different bind params
Same counting bug, different cause. The detector must compare normalized query shapes, not log lines.

## Why it happens

Detectors count query shapes without workload context. To a shape-counter, a bulk import issuing 50,000 similar INSERTs looks identical to an endpoint issuing 50,000 queries per request. The difference is entirely in the surrounding trace: one is a single background job doing its job, the other is a hot path melting under load. Without job tags in the telemetry, the agent cannot see the difference, so it treats both as emergencies. Context is what turns a count into a diagnosis.

## Edge cases

- A job that serves user-facing latency (a synchronous export endpoint that runs a bulk query) is both. Keep it under the request rule and fix the query.
- Jobs that fan out per-tenant (one job run per customer) can hide a real per-tenant N+1 inside the bulk pattern. The bulk rule should still flag per-tenant duplication above its higher threshold.
- If the job framework does not propagate trace tags to the database layer, step 4's `application_name` approach works without any tracing SDK.
- Do not silence the job rule entirely. Imports do regress, and a duration-regression alert on the bulk rule catches what the N+1 rule was accidentally catching.
- When the same codebase serves both requests and jobs, the exclusion must be on the trace tags, not on the code path. Tag at runtime, filter at alert time.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_47_rL_rvmU1dzn54OcduTA
