Whetstone.
EvaluationCode-based evals, and the perform_eval seam
Module 2, Lesson 417 min

Code-based evals, and the perform_eval seam

If you learn one thing in this course, learn the first half of this lesson. It is a two-line fact sitting exactly on the seam between the two things the exam cares most about, which makes it about as exam-shaped as a fact can get.

The two signatures

A code-based evaluator written inline in the LangSmith UI is a function with a fixed name. Where you attach it decides what it receives.

Attached to a dataset: offline
def perform_eval(run, example):
    # `example` carries the reference output, because you wrote one down
    ...
Attached to a tracing project: online
def perform_eval(run):
    # no `example`, because production users do not supply correct answers
    ...

One parameter, present offline, absent online. Not style, not versioning: structure. Offline the platform can hand you a curated example with a reference output attached. Online a real person asked a real question and left, so there is no correct-answer object anywhere in the system and nothing to pass as a second argument.

Every time you feel yourself reaching for rote memory here, re-derive it from does a reference output exist in this situation. You will never get it backwards that way.

The four rules around it

The name is mandatory. The function must be called perform_eval. Not a convention the platform politely prefers; it is how your code gets found and invoked.

Both languages work. Python and TypeScript are both supported and the required name is identical in both. For a TypeScript developer this is the best news in the domain, because the SDK side is nowhere near this even-handed.

No internet access. No fetching ground truth from your own API, no third-party validation service, no call out to another model provider. Everything it scores comes from what it was handed.

It is written inline in the UI. Not code you import from your repo. Which has a practical consequence: it is not in version control, so treat it as configuration to document elsewhere, not as source.

The other surface, and the TypeScript tax

SDK evaluators are a different thing wearing the same taxonomy label. You write them in your repo and pass them into evaluate(). Different surface, different contract, no name requirement.

Python: three positional arguments
def correctness(inputs, outputs, reference_outputs):
    return {"key": "correct", "score": outputs["answer"] == reference_outputs["answer"]}
TypeScript: one object
const correctness = ({ inputs, outputs, referenceOutputs }) => ({
  key: "correct",
  score: outputs.answer === referenceOutputs.answer,
});

Same three pieces of information, two delivery mechanisms, plus a snake_case to camelCase rename on the way. If a question shows three positional parameters, you are looking at Python, and sometimes the question hinges on nothing more.

The rest of the divergences are boring and cheap to be asked about. Repetitions is num_repetitions versus numRepetitions. data is optional in Python, required in TypeScript. aevaluate() is Python-only. Summary evaluators, which score the experiment as a whole rather than row by row, may return a bare primitive in Python only; TypeScript wants the structured return. Unit testing entry points differ: @pytest.mark.langsmith against import * as ls from "langsmith/vitest".

How to use this in the exam

The practical skill is not memorising the table, it is noticing that a language difference could exist at all and thinking to check the docs toggle. The failure mode is confidently reading a Python snippet, mentally translating it, and answering about a TypeScript API that was never designed to mirror it.

When a question is language-specific, the language is load-bearing. Read it as part of the question, not as decoration.

Practice

Try it yourself

Recall

Both signatures, cold

The single highest-yield card in this course. Write both out before checking.

Give the exact signature of an in-UI code evaluator attached to a dataset, and the exact signature of one attached to a tracing project. Then say why they differ.

Reveal answer

Attached to a dataset, which is offline, it is perform_eval(run, example). Attached to a tracing project, which is online, it is perform_eval(run), with no example parameter. They differ because a production trace has no reference output: nobody wrote down the correct answer for a live user's request, so there is no example object to pass. The function name is perform_eval in both cases and in both supported languages.

Quiz

Writing an online code evaluator

You are adding a code-based evaluator in the LangSmith UI, attached to a tracing project, to flag production responses that are not valid JSON.

Which first line do you write?
# ???
    try:
        json.loads(run.outputs["answer"])
        return {"key": "valid_json", "score": 1}
    except Exception:
        return {"key": "valid_json", "score": 0}
  1. Adef evaluate(run, example):
  2. Bdef perform_eval(run, example):
  3. Cdef perform_eval(run):
  4. Ddef perform_eval(example):
Show answer

Correct answer: C — def perform_eval(run):

Tracing project means online, online means no reference output, so the signature is perform_eval(run) with a single parameter. Option 1 is the distractor and it is a good one, because it has the correct function name and is exactly what you would write for the offline case, so it looks right to anyone who learned half the rule. Option 0 fails on the name: the function must be called perform_eval, not evaluate. Option 3 is nonsense in both worlds, since the run is the thing being scored and is never the parameter you drop.

Quiz

Constraints on the in-UI code evaluator

Which of these is NOT true of a code-based evaluator written in the LangSmith UI?

  1. AIt can call an external API to look up ground truth
  2. BIt can be written in Python or in TypeScript
  3. CThe function must be named perform_eval
  4. DIt is written inline in the UI rather than imported from your repo
Show answer

Correct answer: A — It can call an external API to look up ground truth

The in-UI code evaluator has no internet access, so calling out to an external API or fetching ground truth from your own service is unavailable. Everything it scores has to be derivable from what it was handed. Option 1 is true and is genuinely good news for a TypeScript developer, since this is one of the few places in this domain where Python has no advantage. Option 2 is true and is a common way to lose a mark: the name is fixed, not conventional.

Quiz

The shape of a TypeScript row-level evaluator

You are writing a row-level evaluator in the TypeScript SDK, not in the UI. What does your function take?

The Python version, for reference
def correctness(inputs, outputs, reference_outputs):
    return {"key": "correct", "score": outputs["answer"] == reference_outputs["answer"]}
  1. AThree positional arguments, inputs, outputs and referenceOutputs
  2. BA single run argument, since the SDK reuses the in-UI evaluator contract
  3. CTwo positional arguments, run and example, mirroring the perform_eval offline form
  4. DA single object argument, destructured as inputs, outputs and referenceOutputs
Show answer

Correct answer: D — A single object argument, destructured as inputs, outputs and referenceOutputs

TypeScript takes one object. Python takes three positional arguments. Option 0 is the distractor and it is the Python shape wearing camelCase, which is exactly the confusion the question is built on, because most LangSmith examples you will have skimmed are Python. Option 1 is the in-UI code evaluator's online shape, a genuinely different surface with a genuinely different contract, and letting the two blur together is the mistake this lesson exists to prevent.

Recall

Two surfaces, not one

A code snippet in an exam question is partly testing whether you can tell which surface you are looking at, so the distinction is worth being able to state before you read the code.

There are two distinct places a code-based evaluator can live. Name them, and give the argument contract of each.

Reveal answer

The in-UI code evaluator, typed inline into LangSmith, must be named perform_eval, and takes (run, example) when attached to a dataset or (run) when attached to a tracing project. The SDK evaluator, written in your own repo and passed into evaluate(), has no name requirement and takes the row's inputs, outputs and reference outputs: three positional arguments in Python, one destructured object in TypeScript. Both are code-based evaluators in the taxonomy and they share no signature, so the identification question is always which surface am I looking at, and the presence of the word run in the parameter list is the strongest tell for the UI one.

Check

Rebuild the divergence table

Write the Python versus TypeScript differences out from memory in a scratch file, then compare.

You should see

Six rows recovered without help, covering repetitions casing, row-level evaluator argument shape, data optionality, aevaluate, summary evaluator return type, and unit test entry point. Missing one is fine. Missing the argument shape is not, since that is the one most likely to appear as a code-reading question.

Sign in to track your progress →