Automations, and closing the loop
Everything else in this module produces information. Automations are the only mechanism that produces an action, which makes them the piece that turns the lifecycle from a diagram into something that runs.
The anatomy
An automation rule has two halves.
A condition: a filter over traces, plus commonly a feedback condition such as a score below a threshold. Feedback here is uniform, which is the useful part. An evaluator score, a human label from an annotation queue and an end-user thumbs-down all arrive as feedback on a run, so one rule can react to any of them.
An action, and there are three worth knowing:
Add to a dataset. The run becomes an example, and a production failure becomes a permanent offline test.
Add to an annotation queue. The run goes in front of a human, who labels it and can then send it onward.
Trigger an online evaluator. The run gets scored, which is how you apply an expensive judge to a narrow slice rather than to everything.
Those three map onto the only three things you sensibly want done with an interesting run: turn it into a test, get a human to judge it, or measure it.
The edge, finally
The Agent Engineering Lifecycle is a cycle because Monitor feeds Test. Here is that edge, concretely:
Your agent produces a bad answer. An online evaluator scores it low. An automation rule sees the low score, adds the run to a dataset. Next time anyone runs the offline suite, that failure is in it.
Neither half closes the loop alone. The evaluator produces a number and stops. The rule needs a signal to fire on. Together they are a feedback edge that runs while you are asleep, and that is the difference between a lifecycle you drew and a lifecycle you have.
The failure mode worth naming
The obvious first rule is add every thumbs-down to the dataset. It is wrong, and it is wrong in a way that takes about three months to become visible.
Users press thumbs-down for reasons that include their own misunderstanding, a slow response, a correct answer they did not like, and a genuine failure. Only the last belongs in a regression suite. And the harvested examples arrive with no reference outputs, because production has none, so a correctness evaluator has nothing to compare against and your dataset quietly becomes a pile of inputs.
The fix is to route through an annotation queue rather than straight to the dataset. A human confirms it was actually a failure and authors the reference output. Slower, and it is the same argument Align Evals makes about hand-labelling: the expensive human step is the one that cannot be faked, which is exactly why it is load-bearing.
Where this leaves you
You now have the whole cycle. Trace so you can see. Analyse so you can find. Harvest into datasets. Score with evaluators, code-based where you can and judges where you must, aligned against humans before you believe them. Run experiments and read them without overclaiming. Ship. Observe with Insights, measure with online evals, and let automations carry the failures back to the start.
Run it once end to end and the ten questions in this exam section stop being facts to memorise and start being places you have been.
Try it yourself
The anatomy of a rule
Two halves, and being able to name the available actions is what most questions on this topic reduce to.
What are the two parts of an automation rule, and what actions can the second part take?
Reveal answer
A condition and an action. The condition is a filter over traces plus, commonly, a feedback condition such as a score falling below a threshold, and feedback here covers evaluator scores, human labels from an annotation queue, and end-user signals from your own application alike. The documented actions are adding the matching run to a dataset, adding it to an annotation queue for human review, and triggering an online evaluator on it. Those three map neatly onto the three things you might want done with an interesting run: turn it into a test, get a human to judge it, or measure it.
Pick the action for the situation
Users are thumbs-downing roughly one response a day, and you want each one looked at by a person before deciding whether it belongs in your regression suite. Which action fits?
Show answer
Correct answer: B — Add to annotation queue, since a human judgement is required first
The word before is doing the work: you want a human verdict prior to the run entering the dataset, and adding to an annotation queue is exactly that action. Option 0 is the tempting one because thumbs-down does feel like sufficient evidence, and it is the mistake that quietly poisons datasets: users press thumbs-down for reasons including their own misunderstanding, a slow response, or a correct answer they did not like, so unfiltered harvesting adds examples whose reference outputs nobody can write. Option 3 is simply false and worth rejecting cleanly.
The Monitor-to-Test edge
Which combination actually implements the lifecycle's feedback edge without a human in the path?
Show answer
Correct answer: D — An online evaluator producing a score, plus an automation rule that adds low-scoring runs to a dataset
It takes both. The evaluator produces the signal and the automation takes the action, and neither half closes anything alone. Option 0 is the distractor to understand rather than merely reject: scoring production feels like the feedback step, but a score sitting in a project changes nothing until something acts on it. If the acting thing is a human noticing a chart, the edge exists on your architecture diagram and not in your week.
The risk of harvesting unattended
Automations are the most leveraged feature in this module and the easiest to point in a bad direction, so the failure mode is worth naming before you build one.
An automation rule adds every run with a thumbs-down straight into your evaluation dataset. What goes wrong over three months, and what is the fix?
Reveal answer
The dataset fills with examples nobody has validated. Users press thumbs-down for reasons including their own misunderstanding, a slow response, a correct answer they disliked, and genuine failure, and only the last of those belongs in a regression suite. Worse, the harvested examples arrive with no reference outputs, since production has none, so a correctness evaluator has nothing to compare against and the dataset silently becomes a pile of inputs. The fix is to route through an annotation queue rather than straight to the dataset, so a human validates the failure and authors the reference output, which is the same reason Align Evals insists on hand-labelling.
Write one rule on paper
For an agent you know, specify a single automation rule end to end.
You have a filter narrow enough that you could predict roughly how often it fires, a stated action, and an answer to what happens to the item after the action, including who authors the reference output if the destination is a dataset. If you cannot say roughly how often it will fire, the filter is too broad, and a rule you cannot predict is a rule you will end up turning off.