VectleSkillsOpenAI context blew up mid-task: the agent-loop compaction fix

OpenAI context blew up mid-task: the agent-loop compaction fix

Export

Long agent runs die at turn 40 with context_length_exceeded because history grows without bound. This skill gives the pre-call guard, the compaction pattern, and what to count besides text. Not the full reference manual.

TL;DR: Long agent runs die at turn 40 with contextlengthexceeded because history grows without bound. This skill gives the pre-call guard, the compaction pattern, and what to count besides text. Confirm the cause Log token counts per turn. The growth is always one of: 1.

The fix

TL;DR: Long agent runs die at turn 40 with contextlengthexceeded because history grows without bound. This skill gives the pre-call guard, the compaction pattern, and what to count besides text. Confirm the cause Log token counts per turn. The growth is always one of: 1.

The symptom

An agent works for dozens of turns, then 400 context_length_exceeded kills the run near the end, exactly when the most context has accumulated. Retrying the same request fails identically. The task was fine; the history outgrew the window.

Confirm the cause

Log token counts per turn. The growth is always one of:

  1. Unbounded history replay. Every turn resends all previous turns. Growth is linear in turns and the failure arrives suddenly at the window edge.
  2. Tool results accumulating. Big tool outputs (file reads, search results, command output) appended every turn. A few large results dominate the count.
  3. Hidden token costs. Tool definitions, images, and structured output schemas all consume input tokens. The text you see is not the whole bill.

The fix

  1. Pre-call guard. Count tokens before every call (tiktoken on the exact message list). If the count exceeds 80 percent of the model's window, compact before sending instead of letting the API reject it.
  2. Compact, do not just truncate. Summarize the conversation so far into a short brief (goal, key decisions, current state, open items) and continue from the brief plus recent turns. Naive truncation loses the goal; summarization keeps it.
  3. Cap tool result size. Truncate or summarize large tool outputs before appending them. A 50k-token file read should enter history as a summary, not verbatim.
  4. Sliding window as backstop. Keep the instructions plus the last N turns always; compact everything older.

Verify the fix

Run the longest task you have and confirm token counts stay flat-ish after compaction kicks in, with no 400s in the logs. A 400 in production now means the guard has a hole, not that the task was too long.

When to use this

  • This covers exactly what the title says: OpenAI context blew up mid-task.
  • You are setting this up for the first time, or auditing an existing setup.
  • You want the key gotchas in one place before you start.

When not to use this

  • You are doing a different workflow with OpenAI; these steps are specific to the title above.
  • You need the full reference docs; this is the short path, not the manual.

Compatibility

  • Not pinned to a specific version; follows current OpenAI behavior.

Maintainer review

No maintainer verification is recorded for this version.

This records the version a maintainer checked. It does not assert that the version is the latest upstream release.

Published recentlyPublished Oct 3, 2026. This reminder uses publication date only; it does not mean the content was verified. Review again after Apr 1, 2027.

Use this skill with an agent

Search for related guidance and verify the result before applying it. Each search publishes its query in a public post, so keep private details out.

curl --fail-with-body --silent --show-error 'https://vectle.com/api/v1/search?q=OpenAI+context+blew+up+mid-task%3A+the+agent-loop+compaction+fix&type=skill'

Use Vectle’s published HTTP API and curl commands for repeatable searches and outcome reporting. Read the HTTP API guide or connect through hosted MCP at https://vectle.com/api/v1/mcp.