## TL;DR
"Connection reset by peer" during CI artifact uploads means the receiving end closed the TCP connection mid-transfer: usually a timeout on an idle connection, an undersized upload hitting a proxy or registry limit, or a flaky network path. Retry with resume or chunked/multipart uploads, raise client and proxy timeouts, and check whether failures cluster on large artifacts.

## Error / query
```text
"connection reset by peer" in CI artifact uploads
```

## Use this skill when
- Artifact, image, or cache uploads fail intermittently with `connection reset by peer` or `ECONNRESET`
- Large artifacts fail while small ones succeed
- Uploads to a registry, S3-compatible store, or artifact server fail from CI but work locally
- Retrying the step usually fixes it

## Not for this skill when
- Downloads fail with resets (mirror the diagnosis, but check the other end's limits)
- The error is `connection refused` (nothing listening; different problem)
- Uploads fail consistently at 100% (auth or permission at completion, not a reset)
- The failure is DNS resolution (name lookup, not an established connection)

## Steps

### Step 1: Check whether failures correlate with artifact size
```bash
ls -lh [artifact-path]
echo "Compare: do failures start above a size threshold (e.g. 100MB, 1GB)?"
```
Expected: if only large artifacts reset, suspect timeouts or size limits on a proxy, load balancer, or the receiving server. Size-correlated failure is the strongest signal.

### Step 2: Test the upload path manually with verbose output
```bash
curl -v --max-time 600 -T [artifact-file] https://example.com/upload-target 2>&1 | grep -E "reset|Connected|HTTP/"
```
Expected: you see where the reset happens (during headers, mid-body, or near completion). A reset right at the start points at a firewall or proxy; mid-body points at timeouts.

### Step 3: Enable multipart/chunked uploads for large artifacts
```bash
echo "Use the storage backend's multipart upload (s3 multipart, registry chunked upload, CI artifact action's built-in chunking)."
aws s3 cp [artifact-file] s3://[bucket]/[key] --expected-size $(stat -c%s [artifact-file])
```
Expected: the upload splits into parts, so a reset only loses one part and the client resumes it. This converts most transient resets from failures into invisible retries.

### Step 4: Raise timeouts on the client and any proxy in the path
```bash
echo "Check: CI runner HTTP client timeout, reverse proxy (nginx client_max_body_size, proxy timeouts), LB idle timeout."
grep -rn "client_max_body_size\|proxy_read_timeout\|proxy_send_timeout" /etc/nginx/ 2>/dev/null
```
Expected: timeouts comfortably above the worst-case upload duration, and body-size limits above the largest artifact. A 60s proxy timeout with a 5-minute upload is a guaranteed reset.

### Step 5: Add bounded retries around the upload step
```bash
for i in 1 2 3; do
  upload-artifact [artifact-file] && break || { echo "attempt $i failed"; sleep $((i * 10)); }
done
```
Expected: transient resets get absorbed by retries with backoff while persistent failures still surface after 3 attempts. Log each attempt so you can see whether retries are masking a growing problem.

## Variant phrasings

### "ECONNRESET uploading to s3 from ci"
Enable multipart (step 3); the AWS CLI does it automatically above the multipart threshold, so check that the threshold config was not raised.

### "docker push connection reset by peer"
Registry or proxy timeout on large layers. Push with fewer parallel layer uploads and check proxy body/timeout limits.

### "artifact upload fails randomly in github actions"
Usually transient network resets. The built-in artifact actions already retry; if yours does not, wrap it (step 5) and check artifact size.

## Why it happens
TCP resets come from the peer: a proxy or server with an idle/absolute timeout closes the connection, a load balancer kills long uploads, or the network path drops. CI runners often sit behind stricter egress proxies than developer machines, which is why the same upload works locally. Large artifacts simply spend more time exposed to every timeout in the path.

## Edge cases and pitfalls
- Retries without multipart re-upload from zero; on a consistently failing large artifact that never converges, fix the timeout first.
- Some registries rate-limit uploads per IP; parallel CI jobs from one NAT IP can trigger resets that look like network flakiness.
- `connection reset` immediately on connect is a firewall, not a timeout; do not raise timeouts for that, check egress rules.
- IPv6/IPv4 path differences: CI may prefer a broken IPv6 route while local uses IPv4; forcing IPv4 is a valid diagnostic.
- Monitor retry rates; a step that retries on every run is a latent outage, not robustness.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_Z1rZbH5_EHynhhIgs5MfOw
