This mirrors LangChain Academy's Agent Observability and Evaluation course, the one the exam guide labels Building Reliable Agents, and it is the sole prep course covering the Test section: 25% of the LCAE, exactly 10 of the 40 questions. It opens where the official course opens, with observability and tracing, because you cannot evaluate what you cannot see. Then evaluation proper: datasets and reference outputs, experiments and the statistics LangSmith deliberately does not give you, code-based evaluators and the perform_eval signature seam, LLM-as-judge and the two separate alignment features everyone conflates, and pairwise comparison. It closes on production: online evals, the Insights Agent, and the automations that turn a monitored failure back into a test case. The exam is semi-open-book against docs.langchain.com and smith.langchain.com, and a LangSmith org is provisioned per candidate, so this course teaches distinctions and clicks rather than signatures you could look up.
Everything in this course is one cycle, and you have never run the whole cycle in one sitting. That is the capstone.
Start in a tracing project with real runs in it. Use the trace surface to find something genuinely wrong: a slow step, a tool call that should not have happened, an answer that reads badly. Do not pick a failure you already know about, because the point is to practise finding one.
Now harvest it. Get twenty runs into a dataset using at least three different routes, deliberately mixing roughly half good and half bad, because a dataset that is all passes teaches an evaluator nothing. Label all twenty by hand in an annotation queue. This is the part nobody wants to do and the part that makes everything downstream real.
Write two evaluators for the same dataset: one code-based, checking something structural you can be certain about, and one LLM-as-judge for the fuzzy property you actually care about. Run Align Evals on the judge against your human labels and read the alignment score. It will probably be worse than you expected, which is the entire point: before that number existed you were shipping a judge on vibes. Iterate the prompt in the Evaluator Playground, re-run alignment, and stop when the score stops moving rather than when it hits a number you like.
Run the experiment twice, once with repetitions at 1 and once higher, and look at what the standard deviation does to your confidence in the first result.
Then push it forward. Attach the code-based evaluator to the tracing project as an online evaluator, and notice that you had to delete the example parameter to make it work. Set an automation rule that adds low-scoring production runs to the dataset without you. Schedule the Insights Agent on the same project.
You are done when a bad production run can travel, unattended, into the dataset that guards the next deploy. If you can explain to somebody else why the example parameter had to go, you can pass this domain.