Whetstone.
Moving Towards ProductionAutomations, and closing the loop
Module 3, Lesson 415 min

Automations, and closing the loop

Everything else in this module produces information. Automations are the only mechanism that produces an action, which makes them the piece that turns the lifecycle from a diagram into something that runs.

The anatomy

An automation rule has two halves.

A condition: a filter over traces, plus commonly a feedback condition such as a score below a threshold. Feedback here is uniform, which is the useful part. An evaluator score, a human label from an annotation queue and an end-user thumbs-down all arrive as feedback on a run, so one rule can react to any of them.

An action, and there are three worth knowing:

Add to a dataset. The run becomes an example, and a production failure becomes a permanent offline test.

Add to an annotation queue. The run goes in front of a human, who labels it and can then send it onward.

Trigger an online evaluator. The run gets scored, which is how you apply an expensive judge to a narrow slice rather than to everything.

Those three map onto the only three things you sensibly want done with an interesting run: turn it into a test, get a human to judge it, or measure it.

The edge, finally

The Agent Engineering Lifecycle is a cycle because Monitor feeds Test. Here is that edge, concretely:

Your agent produces a bad answer. An online evaluator scores it low. An automation rule sees the low score, adds the run to a dataset. Next time anyone runs the offline suite, that failure is in it.

Neither half closes the loop alone. The evaluator produces a number and stops. The rule needs a signal to fire on. Together they are a feedback edge that runs while you are asleep, and that is the difference between a lifecycle you drew and a lifecycle you have.

The failure mode worth naming

The obvious first rule is add every thumbs-down to the dataset. It is wrong, and it is wrong in a way that takes about three months to become visible.

Users press thumbs-down for reasons that include their own misunderstanding, a slow response, a correct answer they did not like, and a genuine failure. Only the last belongs in a regression suite. And the harvested examples arrive with no reference outputs, because production has none, so a correctness evaluator has nothing to compare against and your dataset quietly becomes a pile of inputs.

The fix is to route through an annotation queue rather than straight to the dataset. A human confirms it was actually a failure and authors the reference output. Slower, and it is the same argument Align Evals makes about hand-labelling: the expensive human step is the one that cannot be faked, which is exactly why it is load-bearing.

Where this leaves you

You now have the whole cycle. Trace so you can see. Analyse so you can find. Harvest into datasets. Score with evaluators, code-based where you can and judges where you must, aligned against humans before you believe them. Run experiments and read them without overclaiming. Ship. Observe with Insights, measure with online evals, and let automations carry the failures back to the start.

Run it once end to end and the ten questions in this exam section stop being facts to memorise and start being places you have been.

Practice

Try it yourself

Recall

The anatomy of a rule

Two halves, and being able to name the available actions is what most questions on this topic reduce to.

What are the two parts of an automation rule, and what actions can the second part take?

Reveal answer

A condition and an action. The condition is a filter over traces plus, commonly, a feedback condition such as a score falling below a threshold, and feedback here covers evaluator scores, human labels from an annotation queue, and end-user signals from your own application alike. The documented actions are adding the matching run to a dataset, adding it to an annotation queue for human review, and triggering an online evaluator on it. Those three map neatly onto the three things you might want done with an interesting run: turn it into a test, get a human to judge it, or measure it.

Quiz

Pick the action for the situation

Users are thumbs-downing roughly one response a day, and you want each one looked at by a person before deciding whether it belongs in your regression suite. Which action fits?

  1. AAdd to dataset, since a thumbs-down is already evidence the answer was wrong
  2. BAdd to annotation queue, since a human judgement is required first
  3. CTrigger an online evaluator, so a judge score decides what enters the suite
  4. DNo automation fits, since human review cannot be triggered by a rule
Show answer

Correct answer: B — Add to annotation queue, since a human judgement is required first

The word before is doing the work: you want a human verdict prior to the run entering the dataset, and adding to an annotation queue is exactly that action. Option 0 is the tempting one because thumbs-down does feel like sufficient evidence, and it is the mistake that quietly poisons datasets: users press thumbs-down for reasons including their own misunderstanding, a slow response, or a correct answer they did not like, so unfiltered harvesting adds examples whose reference outputs nobody can write. Option 3 is simply false and worth rejecting cleanly.

Quiz

The Monitor-to-Test edge

Which combination actually implements the lifecycle's feedback edge without a human in the path?

  1. AAn online evaluator alone, since it scores production runs continuously without a human
  2. BThe Insights Agent alone, since its scheduled runs surface categories of failure unattended
  3. CExtended data retention plus a monitoring dashboard, since matching traces stay available to review
  4. DAn online evaluator producing a score, plus an automation rule that adds low-scoring runs to a dataset
Show answer

Correct answer: D — An online evaluator producing a score, plus an automation rule that adds low-scoring runs to a dataset

It takes both. The evaluator produces the signal and the automation takes the action, and neither half closes anything alone. Option 0 is the distractor to understand rather than merely reject: scoring production feels like the feedback step, but a score sitting in a project changes nothing until something acts on it. If the acting thing is a human noticing a chart, the edge exists on your architecture diagram and not in your week.

Recall

The risk of harvesting unattended

Automations are the most leveraged feature in this module and the easiest to point in a bad direction, so the failure mode is worth naming before you build one.

An automation rule adds every run with a thumbs-down straight into your evaluation dataset. What goes wrong over three months, and what is the fix?

Reveal answer

The dataset fills with examples nobody has validated. Users press thumbs-down for reasons including their own misunderstanding, a slow response, a correct answer they disliked, and genuine failure, and only the last of those belongs in a regression suite. Worse, the harvested examples arrive with no reference outputs, since production has none, so a correctness evaluator has nothing to compare against and the dataset silently becomes a pile of inputs. The fix is to route through an annotation queue rather than straight to the dataset, so a human validates the failure and authors the reference output, which is the same reason Align Evals insists on hand-labelling.

Check

Write one rule on paper

For an agent you know, specify a single automation rule end to end.

You should see

You have a filter narrow enough that you could predict roughly how often it fires, a stated action, and an answer to what happens to the item after the action, including who authors the reference output if the destination is a dataset. If you cannot say roughly how often it will fire, the filter is too broad, and a rule you cannot predict is a rule you will end up turning off.

Sign in to track your progress →