Whetstone.
Production-Ready AgentManaging long conversations
Module 3, Lesson 316 min

Managing long conversations

A four-hour agent run fills its context window. There are two levers for that, they drain different reservoirs, and picking the wrong one produces an agent that compresses diligently and still runs out of room.

Lever one: compress the conversation

SummarizationMiddleware with an explicit trigger
from langchain.agents.middleware import SummarizationMiddleware

SummarizationMiddleware(
    model="gpt-5.4-mini",
    trigger=("tokens", 4000),
    keep=("messages", 20),
)

Two thresholds, both expressed as a ContextSize tuple, and the tuple has exactly three forms:

  • ("messages", 50), a message count.
  • ("tokens", 3000), an absolute token figure.
  • ("fraction", 0.8), a proportion of the model’s maximum input tokens, read from the model profile you met in module 1.

trigger says when to fire. keep says how much recent history survives uncompressed.

About the number 85

You will see 0.85 and 0.10 attached to summarization, and they are real. They are the deep agents preconfiguration: trigger=("fraction", 0.85) and keep=("fraction", 0.10), meaning fire when the context is 85 percent full and keep the most recent ten percent.

They are not the plain SummarizationMiddleware defaults, and the distinction matters in both directions. Assume they are universal and you ship an agent that never compresses. Assume they do not exist and you miss where the figures in the prep material come from.

The 170,000 tokens and 6 messages pair is real too, and it now has a precise home. Deep agents compute their summarization settings from the model, and pick one of two branches: if the model profile exposes max_input_tokens they use the fractions above, and if it does not they use trigger=("tokens", 170000) and keep=("messages", 6) instead. A fraction needs a denominator, so where there is no max_input_tokens to take a fraction of, fixed counts are the only option.

Hold them as alternatives, not as percentages with a safety net underneath. Exactly one branch applies to any given model. And none of the four numbers are SummarizationMiddleware defaults; trim_tokens_to_summarize at 4000 and keep at ("messages", 20) are, and trigger still has none.

Lever two: clear the tool results

ContextEditingMiddleware, defaults shown
from langchain.agents.middleware import ContextEditingMiddleware

ContextEditingMiddleware()          # edits=[ClearToolUsesEdit()]

ClearToolUsesEdit triggers at 100,000 tokens, keeps the 3 most recent tool results, and replaces the older ones with a placeholder, [cleared] by default. It can also clear tool inputs, which it does not do by default, and exclude named tools from clearing entirely.

Which lever, and why often both

Summarization drains the conversation. Context editing drains the tool results. An agent that has held a short conversation while making four hundred tool calls has a tool-result problem, and summarizing the conversation will barely move the number.

The practical tuning rule, once both are on: if the agent loses the plot immediately after a compression, the keep value is too aggressive, not the trigger. Raising the trigger only delays the identical failure, because the messages it was mid-reasoning about are still the ones being paraphrased away.

Practice

Try it yourself

Quiz

What this middleware does out of the box

A long-running agent is configured like this and never compresses anything.

Summarization, as written
agent = create_agent(
    model, tools,
    middleware=[SummarizationMiddleware(model="gpt-5.4-mini")],
)
  1. AIt compresses at 85 percent of the model's input limit, the built-in default
  2. BIt compresses once the history passes 20 messages, the keep default
  3. CIt never compresses, because trigger has no default and nothing was set
  4. DIt compresses at 170,000 tokens, the documented no-profile fallback
Show answer

Correct answer: C — It never compresses, because trigger has no default and nothing was set

trigger has no default on the current constructor, so with nothing configured there is no threshold and no compression happens. The first option is the most dangerous wrong answer because 85 percent is a real number in this ecosystem, it is simply a deep agents preconfiguration rather than a plain SummarizationMiddleware default. keep does have a default, of twenty messages, which is what makes the second option feel symmetric and wrong.

Recall

The ContextSize tuple

One small type expresses every threshold in this lesson, which is why learning its three forms is better value than learning any individual number.

What are the three forms of a ContextSize tuple, and what is each measured against?

Reveal answer

("messages", N) counts messages. ("tokens", N) counts tokens as an absolute figure. ("fraction", F) is a proportion of the model's maximum input tokens, read from the model profile. The fraction form is the one with a dependency: it needs profile data carrying max_input_tokens, and where that is unavailable you supply a custom profile yourself. Both trigger and keep take a ContextSize, which is why the same three forms show up twice.

Quiz

Where 85 percent actually comes from

You have seen the figures 0.85 and 0.10 attached to summarization. What are they?

  1. AThe deep agents preconfiguration, trigger fraction 0.85 and keep fraction 0.10
  2. BThe constructor defaults on SummarizationMiddleware in every configuration
  3. CThe hard limits above which the middleware refuses to run at all
  4. DPercentages of the message count rather than of the token budget
Show answer

Correct answer: A — The deep agents preconfiguration, trigger fraction 0.85 and keep fraction 0.10

Those are the values deep agents preconfigures: trigger at a fraction of 0.85 and keep at a fraction of 0.10. Reading them as universal defaults is the natural error and it is worth being precise about, because a plain SummarizationMiddleware with no trigger set does nothing at all. The last option is a real trap too: fraction is measured against max input tokens, not against message count.

Quiz

Diagnosing a bad compression

An agent starts losing the thread immediately after every compression, forgetting what it was mid-way through doing. Which knob do you move?

  1. ARaise the trigger, so compression happens later in the run
  2. BLower the trigger, so compression happens more often and stays smaller
  3. CIncrease keep, so more recent messages survive intact
  4. DSwitch to a larger summarization model, which paraphrases better
Show answer

Correct answer: C — Increase keep, so more recent messages survive intact

Losing the immediate thread is a symptom of too little recent context surviving, which is the keep value. Raising the trigger is the tempting move because it makes compression rarer and therefore feels safer, but it only delays the identical failure to a later point in the run; the messages the model was mid-reasoning about still get paraphrased away when it finally fires.

Quiz

The other lever

ContextEditingMiddleware defaults to a ClearToolUsesEdit. What does that do, and at what threshold?

  1. ADeletes the oldest messages once the history exceeds 3 messages
  2. BClears older tool results at 100,000 tokens, keeping the 3 most recent
  3. CClears every tool result unconditionally at the end of each turn
  4. DRewrites tool results into a short summary once the window is 85 percent full
Show answer

Correct answer: B — Clears older tool results at 100,000 tokens, keeping the 3 most recent

ClearToolUsesEdit triggers at 100,000 tokens by default, keeps the 3 most recent tool results, and replaces the rest with a placeholder. The first option misreads the keep value of 3 as a message count rather than a count of retained tool results, which is exactly the kind of unit slip that turns a nearly-right answer into a wrong one.

Check

Two levers, two reservoirs

In one sentence each, say what summarization drains and what context editing drains, then say why a long-running agent often needs both.

You should see

Summarization compresses the conversation itself; context editing clears tool results. Your reason for needing both is that they empty different reservoirs, so fixing one does nothing for the other.

Sign in to track your progress →