# This model's maximum context length is 8192 tokens, however you requested 9175 tokens

## TL;DR
Your request is bigger than the model's context window. When this hits during document embedding, the fix is to split oversized texts into token-bounded chunks before sending them, and to retry only the failed batches. Trimming the input is not optional; the request will never succeed as-is.

## The error
```
This model's maximum context length is 8192 tokens, however you requested 9175 tokens (9175 in the messages, 0 in the completion). Please reduce the length of the messages or completion.
```

## Fix it
1. Count tokens before you send. Use tiktoken with the model's encoding and measure every text. Expected: you know exactly which texts exceed the limit.
2. Split oversized texts into chunks that fit, leaving headroom for the completion. A common pattern is chunks of a few hundred tokens with small overlaps. Expected: no chunk exceeds the model's limit.
3. Re-run the failed batch only. Keep a record of which document IDs failed and re-embed just those. Expected: the retry succeeds because every request now fits.
4. For chat calls, trim history or use a model with a bigger window. Expected: the prompt fits and the call returns 200.
5. If single documents legitimately exceed the window, chunk them at ingestion time and store chunk-level embeddings. Expected: search still works, on smaller pieces.

## When to use this
- An agent sees the context-length 400 error on Azure OpenAI during embedding or chat.
- A search-ingestion pipeline dies on one huge document.

## When NOT to use this
- 429 rate-limit errors. Those are quota, not size; retrying blindly is right there and wrong here.
- "Unsupported parameter" errors. Those are API shape problems, not size problems.

## Compatibility
- Azure OpenAI Service, all embedding and chat models. Token limits vary by model; the fix pattern is the same.

### Variant phrasings
- "context_length_exceeded"
- "This model's maximum context length is 128000 tokens"
- "Please reduce the length of the messages"

## Root cause
The limit is per request, not per document. One oversized text in a batch fails the whole batch, which is why ingestion pipelines see this as flaky: most batches pass, then one document kills a batch. The Azure-Samples fix for this exact error was to chunk oversized texts into token-bounded pieces before embedding.

## Edge cases
- Token counts differ between the tokenizer you use locally and the model's. Measure with the model's own encoding.
- Overlap between chunks helps retrieval quality but costs tokens. Keep overlap small.
