AI Models

Managing the Context Window in Long-Running Conversations

Practical strategies for truncating, summarizing, and structuring conversation history to keep LLM responses coherent across hundreds of turns.

Mohammed Saqib9 min read
Close-up of a whiteboard with colorful sticky notes for task organization and planning.
Photo by cottonbro studio on Pexels · Pexels License

Every user message, tool call result, and assistant response adds tokens to the conversation; a single function call with a large JSON response can consume thousands. When the prompt exceeds the model’s context limit (or the cheaper cache limit), earlier turns are silently dropped, breaking references to earlier decisions. The failure mode is subtle: the agent doesn’t error, it just forgets what it decided twenty turns ago, often producing plausible but wrong answers.

Why the context window fills up and what breaks

The math is simple but easy to ignore during prototyping. A 32K context window sounds generous until you run a loop where each turn includes a system prompt (2K tokens), a user message (300 tokens), an assistant response with a tool call (500 tokens), and a tool response (anywhere from 500 to 5K tokens if the API returns a large payload). After ten turns you’re already at 15–20K tokens. After thirty, you’re either paying for a larger context or silently losing the oldest messages.

LLM providers handle overflow differently. OpenAI truncates from the middle of the conversation when you exceed the model’s maximum context. Anthropic’s Claude drops earlier messages when the prompt exceeds its context limit. In both cases the behaviour is deterministic at the API level, but your application logic rarely accounts for it. The result is an agent that appears to work for short sessions and then starts contradicting itself: it approves a refund, then later asks for the refund status as if it never happened. It sets a user preference, then ignores it. Because the LLM doesn’t raise an error, you only notice when a user complains.

The root cause is that we treat the conversation history as an append-only log. It’s not. It’s a bounded buffer, and you, the developer, are responsible for managing that boundary. The model will not help you—it will happily generate a coherent response based on whatever context remains, even if that context is incomplete.

Sliding window truncation with token budget

The simplest fix is to reserve a fixed token budget for history and discard the oldest turns when that budget is exceeded. I usually allocate 8K out of a 32K context for conversation history, leaving the rest for system prompt, current user input, and tool responses. The budget is a hard ceiling: once the history exceeds it, you remove the oldest assistant–user pair until the total fits.

Always keep three things in the window:

  • The system prompt (never truncated).
  • The most recent user message.
  • At least one prior assistant response that contains pending tool call IDs (so the model can still reference in-flight operations).

Implement a token counter that runs per turn. Use tiktoken (or the equivalent for your model) so you know exactly how many tokens each message costs before appending it to the prompt.

import { Tiktoken, encodingForModel } from 'js-tiktoken';
 
const enc = encodingForModel('gpt-4');
 
function countTokens(text: string): number {
  return enc.encode(text).length;
}
 
function truncateHistory(
  history: Array<{role: string; content: string}>,
  budget: number
): Array<{role: string; content: string}> {
  // Always keep system message and last user message.
  const system = history.find(m => m.role === 'system');
  const lastUserIndex = history.length - 1 - [...history].reverse().findIndex(m => m.role === 'user');
  const lastUser = history[lastUserIndex];
  
  // Build a candidate history excluding system and last user.
  const middle = history.slice(1, lastUserIndex);
  let totalTokens = countTokens(system?.content ?? '') + countTokens(lastUser.content);
  
  // Keep most recent messages first.
  const kept: typeof history = [];
  for (let i = middle.length - 1; i >= 0; i--) {
    const msg = middle[i];
    const tokens = countTokens(msg.content);
    if (totalTokens + tokens > budget) break;
    kept.unshift(msg);
    totalTokens += tokens;
  }
  
  return [system!, ...kept, lastUser];
}

This approach works well for stateless tool-use loops where the model doesn’t need to recall facts from early turns. The trade-off is that you lose history entirely once it scrolls out of the window. If your agent needs to remember user preferences or decisions made twenty turns ago, truncation alone will fail.

Summarization as a compression strategy

When the window nears its budget, call the LLM to summarise the conversation so far into a condensed block (200–500 tokens) and replace the history with it. This is more expensive than truncation—each summarisation call costs tokens and latency—but it preserves the gist of the conversation.

The summarisation prompt should ask for key facts, decisions made, and unresolved state. For example: “User requested refund for order #1234, admin approved it, payment gateway returned error 500.” Avoid writing a narrative; the model will compress better if you ask for bullet points or a structured JSON object.

async function summarizeHistory(history: Array<{role: string; content: string}>): Promise<string> {
  const prompt = `Summarize the conversation below. Include all facts, decisions, and unresolved issues. Keep the summary under 400 tokens. Use bullet points.
 
${history.map(m => `${m.role}: ${m.content}`).join('\n')}`;
 
  const response = await callLLM(prompt);
  return response; // typically 200–500 tokens
}

Downside: summarisation calls cost extra tokens and latency; the summary may omit details the user later needs. I reserve summarisation for when truncation would drop something critical—for example, when the history contains a multi-step decision tree that the model must reference later. A good heuristic: if the average turn consumes more than 2K tokens and the conversation exceeds 20 turns, summarise once, then continue appending new turns to the summary.

Structured memory: separating ephemeral from persistent state

Instead of relying on the LLM to infer state from raw history, keep a separate notes object that tracks facts that must survive truncation. This is a small JSON block, always in context, that you update explicitly after each turn. It holds user preferences, confirmed data, error states—anything the agent should not forget.

The notes object has its own token budget (e.g., 1K tokens). If it exceeds that, fall back to summarising the notes content. The key is that you control what goes into notes; the LLM doesn’t write to it unless you call a tool that updates it.

interface AgentNotes {
  userPreferences: Record<string, string>;
  confirmedEntities: string[];
  pendingActions: Array<{id: string; status: string}>;
  errorStates: Array<{code: string; message: string}>;
}
 
function updateNotes(notes: AgentNotes, turn: {role: string; content: string}): AgentNotes {
  if (turn.role === 'tool' && turn.content.includes('error')) {
    notes.errorStates.push({code: '500', message: 'Payment gateway timeout'});
  }
  // ... other rules
  return notes;
}

This pattern is especially useful when you integrate with external APIs. For example, after a user confirms their shipping address, write it into notes. The next turn can read it from notes instead of relying on the conversation history. Without this, a single truncation could erase the address, and the model would ask the user again.

Tool call result pruning and deduplication

Many tool calls return the same data repeatedly—for instance, querying the same order status every few turns. Cache the result and only include it in context if the LLM explicitly asks for a refresh. You can detect a duplicate by comparing the tool call parameters; if the same endpoint with the same arguments was called within the last N turns, skip the response and insert a placeholder: “[cached result from turn 12]”.

If a tool call returns an error or an empty array, replace the full response body with a one-line summary: “GET /orders returned 500”. There’s no need to include the stack trace or the full JSON. Strip redundant fields: if the schema says the response includes a full user object but the LLM only needs the ID, trim the payload server-side before returning it to the conversation.

function pruneToolResponse(response: any, toolName: string): string {
  if (response.error) return `${toolName} returned error ${response.error.code}`;
  if (Array.isArray(response) && response.length === 0) return `${toolName} returned empty array`;
  // Strip user object down to ID if not needed
  if (response.user) return JSON.stringify({ userId: response.user.id });
  return JSON.stringify(response);
}

This practice alone can cut your token consumption by 30–50% in heavy tool-use applications.

Handling multi-turn tool call chains

When the LLM calls tool A, then tool B using A’s output, the chain must remain intact across truncation. If the oldest part of the chain is dropped, the model may try to re-invoke tool A, wasting a call and adding latency. Tag each tool call with a chain ID and preserve the last N complete chains.

A chain starts when the LLM calls a tool and ends when the LLM produces a final response that doesn’t require further tool calls. During chain execution, keep all intermediate messages—tool calls, tool responses, assistant messages that contain reasoning. Once the chain is complete, you can summarise the intermediate outputs: “Tool A returned 3 results, tool B used result #2 to create a ticket.”

Test the failure mode explicitly: what happens if the chain is broken and the LLM re-invokes tool A? In my experience, the model often re-issues the call with the same parameters, producing a duplicate side effect (e.g., creating a second ticket). To prevent this, implement idempotency keys on your tools and detect duplicate calls by comparing the call parameters within a short time window.

Measuring and monitoring context pressure

You cannot manage what you do not measure. Log per-turn token counts for system prompt, history, notes, current user input, and each tool response. Surface a context pressure metric: (used / limit) as a percentage. Alert when pressure exceeds 80% so you can test whether truncation or summarisation is working before the user hits a silent failure.

The table below compares the three main strategies on dimensions that matter in production.

Strategy Token cost per turn Latency impact Fidelity for long-term facts Implementation complexity
Sliding window truncation Zero None Poor – facts are dropped Low
Summarisation Moderate (1–2K tokens per summary) 1–2 additional LLM calls per summary Good – gist preserved Medium
Structured memory Low (notes update) None Excellent – explicit persistence Medium–High

Benchmark your own application, not generic numbers. A 32K context behaves differently when your average turn consumes 2K tokens vs 8K tokens. Instrument your agent loop early and watch the trend over real user sessions.

For deeper reading on token counting pitfalls, see Token counting is wrong: measure actual LLM cost. If you are building an agent loop that makes multiple tool calls, the patterns described in Building an agent loop: tool calls, retries, failure modes are directly applicable.

External documentation: OpenAI’s model overview covers context limits per model (https://platform.openai.com/docs/models). Anthropic’s documentation on context windows explains how messages are dropped (https://docs.anthropic.com/en/docs/build-with-claude/context-windows). Tiktoken’s GitHub repository (https://github.com/openai/tiktoken) provides tokenizer implementations for multiple model families.

Key takeaways

  • Context windows are bounded buffers, not append-only logs. You must actively manage what stays in the prompt.
  • Sliding window truncation is cheap but loses facts. Use it only when the agent does not need long-term memory.
  • Summarisation preserves the gist at a cost. Reserve it for conversations where dropping a single turn would break the agent’s reasoning.
  • Structured memory (a small, always-present notes object) is the most reliable way to keep critical facts across many turns.
  • Prune tool responses aggressively—cache duplicates, truncate errors, strip fields the model does not need.
  • Monitor context pressure per turn. Alert at 80% usage so you can adjust your strategy before the model forgets.

Frequently asked questions

Should I always use summarization instead of truncation?
No. Summarization adds latency and cost per call, and the summary may lose details the user later needs. Use truncation as the default and switch to summarization only when the history contains critical decisions or state that would be lost by dropping old turns.
How do I handle the system prompt if it changes mid-conversation?
Treat the system prompt as part of the token budget. If you need to change it (e.g. user switches to a different task), replace it in full and re-add any persistent memory notes. Never append a second system message; most models only respect the last one.
What happens if the model's context limit is hit during a tool call response?
The tool response is truncated mid-stream or dropped entirely, which can leave the LLM with incomplete data. Always check the token count of the tool response before appending it; if it would exceed the budget, either trim the response server-side or include a placeholder and tell the LLM the data was too large.
#llm#context-window#agents#conversation-management
Share

Keep reading