## TL;DR
Prometheus opens many files: one per chunk plus WAL segments, and the count grows with series count and retention. "Too many open files" means the process ulimit is below what the workload needs. Raise the nofile limit for the Prometheus process (systemd unit, container spec), and check whether runaway series growth is inflating the file count beyond the healthy baseline.

## The query
```text
Prometheus "too many open files": ulimit and storage debugging
```

## Use this when
- Prometheus logs too-many-open-files errors
- Queries fail with file-related errors
- Compaction or WAL replay struggles
- After series count growth

## Not for when
- General Linux ulimit tuning for other services
- Disk space issues (different error)
- Query performance (different topic)

## Steps

### Step 1: Check the current open file count vs limit
Compare the process's open file descriptors against its nofile limit. If usage is at the limit, the limit is the problem; if usage is far below, something else is wrong.
Expected output: usage vs limit quantified.

### Step 2: Raise the nofile limit properly
Set the limit in the process's supervisor (systemd LimitNOFILE, container ulimits, or the shell for manual runs). Raising the system-wide default without setting it for the Prometheus process changes nothing.
Expected output: the Prometheus process running with a limit well above its needs.

### Step 3: Check for file descriptor leaks
If open count grows unboundedly over days, something leaks: usually a storage bug or a stuck compaction. A healthy Prometheus has a stable file count proportional to series and retention.
Expected output: stable file count, or the leak identified.

### Step 4: Correlate with series growth
More series means more chunk files. If the file count jumped with a cardinality explosion, fix the cardinality (drop the bad labels) rather than just raising limits forever.
Expected output: file growth explained by series growth or ruled out.

### Step 5: Verify after the change survives restarts
Confirm the raised limit persists across restarts and redeploys: check the running process's limits, not just the config file. Limits set in the wrong place silently revert.
Expected output: the limit correct on the running process after a restart.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_BYY-RwvI6UpC3R94afuaWA
