Token and cost breakdowns, and the greedy pricing walk
Cost is not a switch you flip. It is a number LangSmith computes, and the computation is more interesting than it looks, because model pricing stopped being one rate per direction a while ago.
Input tokens now split into types. A cache read is not priced like fresh input. Reasoning tokens are not priced like output. So the total has to reconcile across several rates on the same run, and the rule for that reconciliation is one sentence you should be able to quote.
The rule, verbatim
Pricing is applied greedily from most specific to least specific token type.
Walk the types in order of specificity, most specific first. At each step, price that type’s tokens at its own rate, then subtract them from the remaining total. Whatever is left over is priced at the general rate.
The documented shape looks like this:
input_cost = 5 * 1e-6 + (20 - 5) * 2e-6Five tokens at the specific rate, fifteen at the general rate, twenty tokens charged, once each.
The model pricing map
The map is what turns a model name into rates, and two properties of it explain most of the weird behaviour people report.
Matching is by regex, not by exact string. One entry covers a family of model names and their dated variants without anyone hand entering every string a vendor will ever ship. That is also why a model name you never explicitly configured can still be priced correctly.
Entries carry activation dates. A price applies from a date. That is what makes a historical cost view honest: a run from March is priced at March’s rates even after the vendor changes them in June.
Put those together and a symptom that looks like a bug becomes obvious. Recent runs are priced, older runs of the same model are not. Nothing is broken. The entry matches the name, and its activation date is later than those runs.
Why this is such good exam material
It is arithmetic with exactly one conceptual trap in it, it is stated in the docs as a formula you can go and look at, and the wrong answer is one a competent engineer produces by reasoning rather than by ignorance. That combination is what question writers are hunting for.
Try it yourself
Do the arithmetic
A run reports 20 input tokens in total. 5 of those belong to a more specific token type priced at 1e-6 per token. The general input rate is 2e-6 per token.
A: input_cost = 20 * 2e-6
B: input_cost = 5 * 1e-6 + 20 * 2e-6
C: input_cost = 5 * 1e-6 + (20 - 5) * 2e-6
D: input_cost = 20 * 1e-6Show answer
Correct answer: C — Line C, specific tokens are priced first and subtracted from the remainder
Greedy means the most specific token type is priced first at its own rate and then subtracted from the total before the next rate applies. Line B is the trap and it is the mistake competent people actually make: it charges those five tokens twice, once at the specific rate and again inside the twenty. Line D applies the specific rate to fifteen tokens that were never in that bucket, and line A ignores specificity altogether.
What greedy is doing in that sentence
The word is doing precise technical work, not hand waving. Being able to state the algorithm is worth more than remembering the example numbers.
LangSmith applies pricing greedily from most to least specific token type. Describe that algorithm in one or two sentences, and name the error the word greedy exists to prevent.
Reveal answer
Walk the token types from most specific to least specific. At each step, price that type's tokens at its own rate and subtract them from the remaining total, so every token is charged exactly once at the most specific rate that applies to it. Whatever is left at the end is priced at the general rate. The error it prevents is double counting: adding a specific rate on top of a general rate that already covered the same tokens, which overstates cost and is the natural wrong assumption.
How a pricing entry is matched
Your model name is slightly different from anything you remember configuring, and a run from two months ago has no cost while a run from yesterday does.
Show answer
Correct answer: A — Entries match by regex on the model name and carry activation dates, so one can match the name yet not cover an older run
Regex matching plus activation dates explains both halves of the symptom in one go: the name matches, the date does not. The nightly recomputation answer is the interesting distractor because plenty of billing systems do backfill, and it sounds like a kindness. It would defeat the point of activation dates, which exist precisely so that historical spend is not silently rewritten when a vendor changes a price.
Why breakdowns exist at all
Modern providers no longer bill one flat rate per direction. What is the most common reason a single run's input tokens split across two rates?
Show answer
Correct answer: B — Part of the input was a cache read, which is priced lower than fresh input
Cached input reads are the everyday reason a breakdown exists, and they are cheaper than fresh input, which is exactly why providers report them as their own token type. The bulk rate option is the tempting one because volume tiering is common in infrastructure pricing generally, but that would reprice the whole request rather than split it, which is not the shape the greedy walk describes. Streaming does not change token pricing.
Price one real run by hand
Take a single run out of any project you have and do the sum on paper before you look at the answer LangSmith computed.
Working from the run's token counts and the applicable rates, you produce a number by hand, then compare it against the number in the cost column, and the two agree. If they do not, the gap is almost always a token type you did not notice was priced separately.