Whetstone.
Models and Structured OutputStructured output, and the chronicler gap
Module 1, Lesson 225 min

Structured output, and the chronicler gap

Structured output is the interesting part. You do not want prose back, you want a typed object that matches a schema. Define the schema in zod, then bind it to the model.

Structured output with withStructuredOutput
import * as z from "zod/v4";

const Decision = z.object({
  messageId: z.string(),
  category: z.enum(["A", "B", "C"]),
  confidence: z.enum(["low", "medium", "high"]),
  reason: z.string().optional(),
});

const structured = model.withStructuredOutput(Decision);
const result = await structured.invoke(messages); // typed as Decision

Under the hood, withStructuredOutput uses the provider’s tool-calling or structured-output mechanism to force the model to return JSON that fits the schema, then parses it for you. That parse-and-validate step is exactly what chronicler hand-rolls today.

withStructuredOutput should still exist in the 1.5.x line; verify the exact surface at build time, since the modern alternative inside an agent loop is toolStrategy, which lesson 2 of module 2 covers.

Here is what the framework gives you for free, set against what chronicler’s requestStructuredJson does by hand:

Map it across: requestStructuredJson’s safeParse maps to withStructuredOutput’s built-in parse. The four-attempt backoff maps, roughly, to LangChain’s model-level maxRetries / retry config. The fence-stripping is mostly unnecessary once you use tool-calling structured output, because the model returns structured content, not a markdown code block. The preservation-bias default is business logic that stays yours either way, no framework replaces a product decision.

Practice

Try it yourself

Quiz

Why fence-stripping mostly disappears

Chronicler's markdown-fence stripping becomes mostly redundant once you use LangChain structured output. Why?

  1. AwithStructuredOutput returns structured data via tool-calling, not a markdown fence
  2. BLangChain quietly disables all markdown formatting globally
  3. CFence-stripping was actually never necessary in the first place
  4. Dzod automatically strips markdown code fences for you
Show answer

Correct answer: A — withStructuredOutput returns structured data via tool-calling, not a markdown fence

Tool-calling / structured-output mode gets JSON back as structured content, not prose wrapped in a markdown fence. The fence-stripping heuristic exists because chronicler's older path could get prose-wrapped JSON back.

Recall

Where the preservation bias lives

Chronicler defaults omitted messageIds to category B. Keep the distinction between what a schema can validate and what it can invent.

Where does the omitted-id-defaults-to-B rule live, and why can't the schema itself express it?

Reveal answer

It lives in your own business logic, applied after the model call. A zod schema can validate the shape of a response the model gives; it cannot invent a default for a message the model never mentioned. That decision is yours to keep regardless of which LLM plumbing you use.

Do

Re-implement a day of chronicler classification

Re-implement one day-chunk of chronicler's A/B/C classification using LangChain structured output, then diff it against the cached, hand-rolled result for that same day.

  • Pick a real transcript day chronicler has already cached under its per-day cache.
  • Write a spike file in ~/dev/lumiere/apps/chronicler (not wired into the pipeline) that runs the same prefiltered messages through model.withStructuredOutput(Decision).
  • Diff your run's categories against the cached @lumiere/openai-json output for that day.
  • For every difference, write down why it differs, don't hand-wave it.
  • Write a short verdict: what requestStructuredJson's 4-attempt / fence-strip / safeParse dance maps to in LangChain, and what it does not cover (the specific retry policy, the preservation-bias default, anything chronicler-specific).
Done whenYour spike matches the cached day's categories for more than 95% of messages, every difference is explained, and you have a written adopt/skip verdict on chronicler's LLM plumbing.
Check

Classification parity

Confirm the spike's output against the cached day.

You should see

Same categories as the cached day for more than 95% of messages, with every difference explained.

Sign in to track your progress →