## TL;DR
Do not optimize the pipeline by feel: pull per-job durations from your CI's API or timing view, separate queue time from run time, and find the critical path (the longest chain of dependent jobs). The slowest job on the critical path is your target; speeding up anything else does not move the total. Queue time often dominates, and that is a runner capacity problem, not a job problem.

## Error / query
```text
how to find the slowest job in a CI pipeline
```

## Use this skill when
- The pipeline is too slow and you need to know where time goes
- Deciding what to parallelize or cache first
- Measuring whether an optimization actually helped
- Queue times feel longer than the work itself

## Not for this skill when
- One test is slow (test-level profiling)
- Runners never provision (infra problem)
- The pipeline is flaky rather than slow (reliability problem)

## Steps

### Step 1: Get per-job durations from the CI API
```bash
gh run view [run-id] --json jobs -q '.jobs[] | "\(.name): queued=\(.startedAt) run"'
# or use the CI timing view / pipeline graph for the same data
```
Expected: each job's queued-at, started-at, and completed-at timestamps. Compute queue time (queued to started) and run time (started to completed) separately; they have different fixes.

### Step 2: Rank jobs by run time on the critical path
```text
List jobs in dependency order. The critical path is the longest
dependency chain; sum run times along each chain. The job with
the largest run time ON the critical path is the target.
```
Expected: one job (or chain) clearly dominates. A 20-minute test job off the critical path matters less than a 5-minute build job everything waits on.

### Step 3: Separate queue time from run time
```bash
# queue_time = startedAt - createdAt ; run_time = completedAt - startedAt
# aggregate over the last 20 runs: p50 and p95 of each
```
Expected: if p95 queue time dominates, add runners or fix labels; if run time dominates, optimize the job (caching, parallelization, smaller scope). Fixing the wrong one wastes the effort.

### Step 4: Re-measure after each change
```text
Record: total pipeline p50/p95, critical-path job, its run time.
Change one thing. Compare the same metrics over the next 20 runs.
```
Expected: the numbers move or they do not. Pipeline speed work without before/after measurement is guessing; the API gives you the data in minutes.

## Variant phrasings

### "ci pipeline slow which job"
Steps 1-2: per-job timings, then critical-path ranking.

### "github actions slow jobs timing"
The run view and job timings (step 1); separate queue from run time (step 3).

## Why it happens
Pipeline time is the critical path plus queueing, and intuition about which job is slow is usually wrong: the loud complaint is the job that fails, while the quiet time sink is an un-cached build step everything depends on. Measurement finds the real bottleneck, and splitting queue vs run time points at capacity vs efficiency fixes.

## Edge cases and pitfalls
- Averages hide the pain; use p50 and p95, since the slow runs are what developers feel.
- Matrix jobs multiply: one slow matrix leg times 20 combinations is the real cost; look at per-leg times, not the matrix total.
- Cache warm-up runs are outliers; exclude the first run after cache invalidation from the baseline.
- Speeding up a job can shift the critical path; re-rank after each optimization instead of assuming the same target.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_NXc9i54S8KF21VtoLQg8Wg
