Whetstone.
ObservationObservability, and the death of assertEquals
Module 1, Lesson 115 min

Observability, and the death of assertEquals

Everything you know about testing assumes the function returns the same thing twice. Agents do not, and pretending otherwise is how teams end up with a suite everyone has learned to re-run rather than read.

The assertion problem

A pure function has one right answer, so expect(f(x)).toBe(y) is a complete test. A model has a distribution of acceptable answers, and the good ones differ from each other. Assert equality and you get red on output that was genuinely fine, every few runs.

The usual response is to weaken the assertion until it stops complaining. Check the output is non-empty. Check it is a string. Check it contains one keyword. Now the test passes reliably and asserts almost nothing, which is the worst of both worlds: it costs CI time, it occupies a slot in your head labelled tested, and it would not catch the agent answering in German.

So the assertion gets replaced by a score. You stop asking whether this is the right string and start asking whether the output has the property you care about, then you watch that score across versions. That scoring function is an evaluator, and most of this course is about evaluators, what you run them against, and how you know whether to trust them.

Why the score alone is not enough

Here is where the official course order earns itself. You run the evaluation, accuracy falls from 0.86 to 0.71, and you now know precisely one thing: something got worse.

Not which step. Not whether the retriever returned nothing, or the model was handed a truncated context, or a tool started timing out and the agent retried it four times and gave up. A score is an alarm. A trace is the diagnosis. You need both, and you need the trace first, because a number with no way to investigate it produces a team that stares at a dashboard and guesses.

What good looks like

A well-observed agent lets you answer, for any single request, what the user asked, what the agent decided, which tools it called with which arguments, what those tools returned, what the model was actually shown, how long each step took, and what it cost.

Notice how many of those are unavailable from the final output alone. The final output is one string at the end of a tree, and almost every interesting failure happens further up.

The reframe worth carrying

Testing an agent is not one activity. It is two, and they need different instruments.

Measurement answers “is this version better than that one”, and its instrument is the evaluator. Investigation answers “why did this specific request go wrong”, and its instrument is the trace. Teams that own only the first ship confidently in the wrong direction. Teams that own only the second debug beautifully and cannot tell whether the product is improving.

Practice

Try it yourself

Recall

Why assertEquals dies on contact with an agent

No looking back. The testing style you have used for twenty years stops working at a specific point, and it is worth being able to say exactly where.

You have a function that takes a support ticket and returns a summary. Why can you not test it with an equality assertion against a known-good string, and what replaces that assertion?

Reveal answer

The model is nondeterministic, so the same input produces different valid strings across runs. An equality assertion fails on outputs that are completely correct, which makes it worse than no test at all because it trains the team to ignore red. What replaces it is an evaluator: a function that scores the output on some property you actually care about, such as whether it mentioned the refund, whether it is under fifty words, or whether it is grounded in the ticket, rather than comparing it byte for byte against one blessed answer. The unit of testing moves from is this exactly right to is this good enough, and is it better or worse than last time.

Quiz

Why a flaky assertion is worse than nothing

An equality assertion against a model output fails perhaps one run in four. Why is that specifically worse than having no test on that behaviour at all?

  1. AIt burns CI minutes that could be spent on faster tests
  2. BIt cannot be retried, so the pipeline blocks permanently
  3. CIt produces no score, so it cannot be compared across versions
  4. DIt teaches the team that a red build carries no information
Show answer

Correct answer: D — It teaches the team that a red build carries no information

A test that cries wolf reliably destroys the signal value of every other test beside it, because people learn to re-run rather than to read. That is a cost paid by the whole suite, not just by the flaky test. Option 2 is the tempting one because it is true and it is the thing this course goes on to fix, but it describes a missing capability rather than active harm. The failure mode being named here is that the flaky assertion makes your existing tests worse.

Recall

Why observation comes before measurement

The official course opens here rather than with evaluation, and the ordering is an argument rather than a preference.

Why does observability have to come before evaluation, rather than the other way round or in parallel?

Reveal answer

Because an evaluation score is an alarm and a trace is the diagnosis. A score tells you a number moved between two versions; it cannot tell you that the retrieval step returned nothing, that the agent called the same tool four times, or that the model saw a truncated prompt. Without traces you can detect that something is wrong and never find out what. There is a practical dependency too: the production runs you eventually harvest into datasets only exist because they were traced, so an untraced application cannot grow its own test suite.

Quiz

What observability buys that a score does not

Your evaluation shows accuracy dropped from 0.86 to 0.71 after a change. Which capability does observability add that the score alone cannot?

  1. AIt tells you whether the drop from 0.86 to 0.71 is statistically significant
  2. BIt tells you which failing examples to add to the dataset next
  3. CIt lets you open a failing request and read every intermediate step
  4. DIt automatically re-runs the failing examples and reports the new score
Show answer

Correct answer: C — It lets you open a failing request and read every intermediate step

Tracing gives you the tree of intermediate steps for a specific request: what each tool returned, what the model was actually shown, where the latency went. That is the only route from a moved number to a cause. Option 0 is a good distractor because significance is exactly what you want at that moment and it is a thing the tooling does not provide anywhere, as a later lesson covers. Option 3 describes automation rules, which are real but are a different feature.

Check

Audit a suite you already own

Take any service you have shipped that calls a model, and look at how it is tested today.

You should see

You can name at least one assertion in it that is either flaky or has been weakened until it asserts almost nothing, such as a check that the output is non-empty. That weakening is the fingerprint of an equality assertion meeting nondeterminism and losing, and the honest replacement is a score you watch over time rather than a boolean you gate on.

Sign in to track your progress →