Whetstone.
Context ManagementSummarization and offload, with the actual numbers
Module 3, Lesson 119 min

Summarization and offload, with the actual numbers

An agent that runs for four hours does not fail because the model got worse. It fails because the context window filled up with tool output nobody needed.

The failure this module exists for

Long-running agents die of their own success. Every tool call returns something, every something gets appended, and the window is finite. The killer is almost never the conversation, it is tool results: a file read, a search hit list, a database dump, an API response with two thousand rows in it.

So the defences are aimed there, and they are mechanical rather than clever. There are two, they drain different reservoirs, and a serious agent needs both.

Offloading: 20,000 tokens

The tool-result offload threshold is 20,000 tokens. Past that, a result does not go into the context. It goes to a file.

That figure is a constructor argument, not folklore: FilesystemMiddleware(tool_token_limit_before_evict=20000). The same constructor carries a second, less-quoted threshold: human_message_token_limit_before_evict=50000, which does the same job for an oversized human message. Two thresholds, two message types, and the larger one is the easier exam distractor precisely because most study notes only mention the first.

Summarization: whose defaults are these

Here is the boundary that the numbers below hang off, and it is the single most misread thing in this module.

That function picks one of two mutually exclusive branches, based on whether the model profile exposes max_input_tokens:

compute_summarization_defaults, both branches
profile exposes max_input_tokens:
  trigger                 ("fraction", 0.85)
  keep                    ("fraction", 0.10)
  truncate_args_settings  ("fraction", 0.85) / ("fraction", 0.10)

no profile info:
  trigger                 ("tokens", 170000)
  keep                    ("messages", 6)
  truncate_args_settings  ("messages", 20) / ("messages", 20)

Read that as alternatives, not as percentages plus fallbacks. Exactly one branch applies to any given model. The fixed figures are not a safety net underneath the fractions; they are what you get instead of them, and the source comments call them deliberately conservative because they are guarding an unknown limit.

It fires at 0.85 and keeps 0.10. Say those as a pair every time, because the exam distractor is the one that reuses 0.85 for both. Consistency feels right and the real design is deliberately lopsided.

What survives is a summary of the older conversation plus the most recent slice kept intact. That shape is deliberate: recent messages are the ones the model is mid-reasoning about, so paraphrasing them would break the thread it is currently pulling on. Old ones are safe to compress because only their conclusions still matter.

Which gives you a real tuning rule: if an agent loses the plot right after a compression, the keep-ratio is too aggressive, not the trigger. Raising the trigger only delays the same failure.

The eighty-five collision, corrected

0.85 really does appear twice, and the tempting explanation is that the two uses measure against different denominators, one the context window and one max_input_tokens. That reading is wrong. Both are ("fraction", 0.85), and a fraction is only ever resolved one way: against the model profile’s max_input_tokens.

The two uses differ in what they do, not in what they measure:

  • truncate_args_settings clips oversized tool-call arguments on messages older than the keep window. Think write_file content and edit_file patches.
  • The summarization trigger compacts the conversation itself.

And the ordering is the point: argument truncation is the cheap pre-pass, and it frequently reclaims enough room that full summarization never has to run at all. One threshold, two escalating responses.

Where it sits in the stack

SummarizationMiddleware is part of the fixed deep agent base stack, after FilesystemMiddleware and the subagent middleware and before PatchToolCallsMiddleware. Its exact ordinal position moves, because the middleware ahead of it is conditional: skills middleware is only present if you passed skills, and subagent middleware only if the agent has inline subagents. Learn it by neighbours, not by number.

One more precision worth carrying: what a deep agent installs is deepagents’ own wrapper around the LangChain middleware, not the plain class. The wrapper adds backend offload of evicted history to /conversation_history/{thread_id}.md, the pre-summarization argument truncation above, a retry on provider context-overflow errors, and non-mutating message state. The plain SummarizationMiddleware drops evicted messages with no recovery path. Defaults and assembly, not capability, one more time.

Practice

Try it yourself

Quiz

The offload threshold

A number question. Semi open book means an examiner can ask this, but only because it is quick to state and easy to confuse with the other figures in this lesson.

  1. A8,000 tokens
  2. B50,000 tokens
  3. C20,000 tokens
  4. D170,000 tokens
Show answer

Correct answer: C — 20,000 tokens

The tool-result offload threshold is 20,000 tokens, set by tool_token_limit_before_evict on FilesystemMiddleware. 50,000 is the genuinely dangerous distractor because it is also a real eviction threshold on the same constructor, just the one for oversized human messages. 170,000 is the other real number in this lesson, from the no-profile summarization branch.

Recall

What survives an offload

The mechanism matters more than the threshold, because it explains why the agent can still work with a result it can no longer see.

When a tool result is large enough to be offloaded, what does the agent actually have in its context afterwards?

Reveal answer

The result is replaced by a file path plus a preview, and the preview is head-and-tail rather than head-only: the first 5 lines, then a marker of the form ... [N lines truncated] ..., then the last 5 lines, all with line numbers. The path is a handle the agent can use later with read_file, grep or glob to pull back only the slice it needs, and the N in the marker tells it how much it is not seeing. Nothing is lost. It is moved out of the window and left addressable, which is why the filesystem tools are defaults rather than an optional extra.

Quiz

Trigger and keep in a deep agent

The two fractions a deep agent configures for summarization. They are not the same number and they are not the same kind of number.

  1. ATrigger at fraction 0.85 of max_input_tokens, keep fraction 0.10
  2. BTrigger at fraction 0.85 of max_input_tokens, keep fraction 0.85
  3. CTrigger at fraction 0.10 of max_input_tokens, keep fraction 0.85
  4. DThese are SummarizationMiddleware's own constructor defaults
Show answer

Correct answer: A — Trigger at fraction 0.85 of max_input_tokens, keep fraction 0.10

A deep agent whose model profile exposes max_input_tokens gets trigger fraction 0.85 and keep fraction 0.10. Say them as a pair every time, because the distractor that reuses 0.85 for both feels internally consistent and is not: one number says when to act, the other says how much survives, and the design is deliberately lopsided. The last option is the structural trap: these come from deepagents' compute_summarization_defaults, not from SummarizationMiddleware, whose own trigger has no default at all.

Quiz

Two eighty-fives, one denominator

0.85 appears twice in a deep agent's summarization config, governing two different mechanisms. What is the relationship?

  1. ATwo different denominators, one the context window and the other max_input_tokens
  2. BOne fraction of max_input_tokens driving both; the cheaper truncation runs first
  3. COne is a fraction and the other a raw token count that happens to look similar
  4. DOnly one 0.85 exists; the other figure is 0.80 and governs truncation
Show answer

Correct answer: B — One fraction of max_input_tokens driving both; the cheaper truncation runs first

Both are ("fraction", 0.85) and both resolve against the same quantity, the model profile's max_input_tokens. The difference is what they do, not what they measure: truncate_args_settings clips large tool-call arguments in older messages, while the summarization trigger compacts the conversation. The first option is the seductive one because two mechanisms sharing a number really does invite the idea that they must be measuring different things, and that reading was in circulation; it is wrong. The argument pass is the cheap pre-step and often reclaims enough context that full summarization never has to run.

Recall

The no-profile branch

What a deep agent does when the model profile carries no max_input_tokens. Two absolute figures, and they look nothing like the fractions, which is what makes them memorable.

What values does a deep agent use for summarization when the model exposes no max_input_tokens, and why does that branch exist at all?

Reveal answer

Trigger becomes ("tokens", 170000) and keep becomes ("messages", 6), with argument truncation dropping to ("messages", 20) for both its trigger and its keep. The branch exists because a fraction is meaningless without a denominator: fraction-based thresholds are computed against the model profile's max_input_tokens, so where no profile exposes that number the middleware cannot use fractions and falls back to fixed counts. These are alternatives to the 0.85 and 0.10 pair, not additions to it; exactly one branch applies to any given model, and the comment in the source calls the fixed figures deliberately conservative to avoid overshooting an unknown limit.

Check

The numbers, cold

Close the lesson and write the numbers down with what each one governs, before scrolling back.

You should see

You wrote 20,000 tokens for tool-result offload and 50,000 for oversized human messages, a 5-line head plus 5-line tail for the preview, fraction 0.85 with keep 0.10 for a deep agent whose model profile exposes max_input_tokens, and 170,000 tokens with 6 messages for one that does not, and you attached each to its own mechanism rather than to a general idea of context management. You also kept the boundary straight, in that these are deepagents' choices and SummarizationMiddleware on its own has no trigger default.

Sign in to track your progress →