## TL;DR
NLTK's sentence tokenizer needs the `punkt` resource downloaded into `nltk_data`, and fresh environments dont have it. Run `nltk.download('punkt')` once (and `punkt_tab` on newer NLTK), ideally in your setup script rather than at request time.

```text
nltk LookupError: punkt tokenizer not found in nltk_data
```

## Use this when
- `sent_tokenize` raises LookupError for punkt
- NLP code works on your laptop but fails on a server or container
- An agent set up tokenization without downloading resources

## Not for this skill when
- The missing resource is `punkt_tab` (newer NLTK split; download that too)
- The download itself fails (thats network or SSL)
- You need a non-English tokenizer (separate resource)

## Steps

1. Download the missing resource:

```python
import nltk
nltk.download('punkt')
nltk.download('punkt_tab')  # needed on NLTK 3.8+
from nltk.tokenize import sent_tokenize
print(sent_tokenize("Hi there. How are you?"))
```
Expected output: the download completes and tokenization returns two sentences. The error message itself suggests this command; it is the fix.

2. Verify where NLTK looks, so the download lands somewhere persistent:

```python
import nltk
print(nltk.data.path)
```
Expected output: the search paths. In containers, the default may be a temp dir that vanishes on restart.

3. Bake the download into the image or setup script:

```bash
python -m nltk.downloader punkt punkt_tab -d /usr/share/nltk_data
```
Expected output: resources land in the shared path. Runtime code then never needs network access for tokenizers.

4. Pin the NLTK version so the resource set stays stable:

```bash
pip freeze | grep -i nltk
```
Expected output: a pinned version like `nltk==3.9.1`. Resource splits happen across versions; pinning documents which set your code needs.

## Variant phrasings

### LookupError suggests the download but it still fails after
The download went to a path not in `nltk.data.path` for the running process. Check step 2 in the same process that tokenizes.

### works in notebook, fails in the deployed app
The notebook downloaded to your home dir; the app runs as a different user. Use the shared path from step 3.

## Why it happens
NLTK ships code and data separately; the tokenizer models download on demand. Developers run the download once on their laptop and forget it, then the code fails everywhere else. The LookupError text is unusually helpful (it prints the exact download call), but people skim past it to the traceback instead.

## Edge cases
- Corporate proxies break the NLTK downloader's HTTPS; set the proxy env vars or vendor the files.
- `nltk.download()` without args opens a GUI picker; always pass the package name in scripts.
- If disk is read-only at runtime, the download cant happen lazily; the bake-in step is mandatory.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_PLfFushex8c8Q8epmehq_g
