Whetstone.
The SpellingLists, dicts, and the comprehension
Module 1 · Lesson 216 min

Lists, dicts, and the comprehension

This is the lesson where Python stops being merely different and starts being noticeably nicer. Four built-in containers, one syntax for transforming them, and a slicing notation you will miss when you go back to TypeScript.

The four containers

One line each
names = ["ada", "grace", "alan"]        # list   -> Array. Mutable, ordered.
point = (3, 4)                          # tuple  -> a fixed positional record. Immutable.
config = {"retries": 3, "debug": True}  # dict   -> Map / object. Ordered since 3.7.
seen = {"ada", "grace"}                 # set    -> Set. Unordered, unique, fast membership.

A few notes that are not obvious from the shapes.

list is Array and behaves the way you expect, except the methods are named differently: append not push, pop with an optional index, extend not concat, and len(x) rather than x.length. That last one is a language-wide pattern: Python puts these operations in functions that dispatch to a protocol, rather than in methods on each type. len, iter, str, repr, hash, abs are all this.

tuple has no TypeScript equivalent worth naming. It is a fixed-length, immutable, positional record. Functions return them constantly, and the caller unpacks them. Crucially it is hashable, so it can be a dict key or a set member, which a list cannot. The comma is what makes a tuple, not the parentheses: x = 1, is a one-element tuple and is a genuinely nasty typo.

dict is a Map, not an object. Keys are any hashable value, not just strings. Order is insertion order and that is a guaranteed language feature since 3.7, not an implementation detail.

set is Set, with real operators: a | b union, a & b intersection, a - b difference. Deduplicating a list is list(set(items)), and it loses order.

Missing keys are loud

Three ways to read a key, and they behave differently
config["timeout"]              # KeyError if absent. Loud, local, immediate.
config.get("timeout")          # None if absent. The JavaScript behaviour, opted into.
config.get("timeout", 30)      # A default. This is the one you want most of the time.

Square-bracket access raising on a missing key is the biggest behavioural gap between a dict and a JavaScript object, and it is Python being right. A missing property that quietly evaluates to undefined gives you a TypeError four call frames later, at a place with no information about what actually went wrong. A KeyError names the key and stops there.

The same shape applies to attributes: obj.missing raises AttributeError, and getattr(obj, "missing", default) is the soft form.

Unpacking is everywhere

Destructuring, Python-flavoured
x, y = point                        # tuple unpacking
first, *rest = names                # star capture. rest is a list.
a, b = b, a                         # swap, with no temp variable

for key, value in config.items():   # unpacking in the loop header
    print(key, value)

def send(url, **options): ...
send("https://x", timeout=5)        # ** collects into a dict
send("https://x", **defaults)       # ...and splats one back out

* and ** do double duty, and which duty depends on where they appear. In a function signature they collect. At a call site or in a literal they spread. That is the whole rule, and it is the same rule as ... in TypeScript wearing two different hats.

Spreading into literals
merged = {**defaults, **overrides}   # object spread
combined = [*a, *b]                  # array spread

The comprehension, which is the one you actually need to read fluently

This is the single most characteristic piece of Python syntax, and it appears in roughly every file you are about to open.

filter then map, in one expression
active_names = [user.name for user in users if user.is_active]

Read it in three pieces, left to right: what I am collecting (user.name), where it comes from (for user in users), which ones qualify (if user.is_active). The output expression coming first is the part that feels backwards for about a day and then feels correct forever, because it puts the interesting part at the front instead of at the end of a .map(...) chain.

The bracket you use decides the type:

Four brackets, four results
[x * 2 for x in nums]            # list
{x * 2 for x in nums}            # set
{k: v for k, v in pairs}         # dict
(x * 2 for x in nums)            # GENERATOR. Lazy. Not a tuple.

That last one is worth sitting with. Parentheses give you a generator expression, which produces values one at a time and can only be consumed once. It is the memory-efficient form, and it is why sum(x.cost for x in items) is idiomatic: the sum consumes the generator as it goes and never builds an intermediate list.

Slicing

Slicing, which is worth the whole lesson on its own
rows[1:3]      # items 1 and 2. Start inclusive, stop exclusive.
rows[:3]       # first three
rows[3:]       # everything from index 3
rows[-1]       # last item
rows[:-1]      # everything except the last
rows[::2]      # every other item
rows[::-1]     # reversed

The bounds clamp rather than raise. rows[0:10] on a three-item list gives you three items and no complaint, while rows[10] raises IndexError. Both behaviours are useful and the asymmetry is deliberate, but it means a slice will never warn you that your length assumption was wrong.

Strings slice identically, because a string is a sequence of characters. There is no separate substring method.

Strings, briefly

f-strings do more than interpolate
name = "ada"
count = 3
ratio = 0.8317

f"hello {name}"              # 'hello ada'
f"{count} items"             # '3 items'
f"{ratio:.1%}"               # '83.2%'   format spec after the colon
f"{count=}"                  # 'count=3' self-documenting, great for logging

The f prefix is mandatory. Without it the braces are literal characters, and a forgotten f is the most common cause of a log line that says Processing {job_id} in production.

Two other prefixes are worth recognising when you see them: b"..." is a bytes literal, which is a genuinely different type from a string and will not silently convert, and r"..." is a raw string, which turns off backslash escapes and exists mostly for regular expressions and Windows paths.

Where this goes

Next up is the short list of Python behaviours that will actively lie to you if you assume they work like TypeScript: truthiness on containers, None handling, equality, and the default-argument trap that is genuinely famous for a reason.

Practice

Try it yourself

Quiz

Reading a key that is not there

You have a dict called config with no timeout key, and you evaluate config["timeout"]. What happens?

  1. AIt raises KeyError, and the safe form that returns None instead is config.get("timeout")
  2. BIt returns None, the way a missing property on a JavaScript object returns undefined
  3. CIt returns an empty string, because a dict coerces a missing value to the empty form of its value type
  4. DIt raises AttributeError, which is the error a dict raises for any key it does not hold
Show answer

Correct answer: A — It raises KeyError, and the safe form that returns None instead is config.get("timeout")

Square-bracket access on a missing key raises KeyError. This is the single biggest behavioural difference from a JavaScript object, where a missing property is quietly undefined, and it is a genuine improvement: the failure is loud and local instead of a None that travels three functions before crashing. .get(key) opts into the JavaScript behaviour and .get(key, default) supplies a fallback. AttributeError is the error for a missing attribute on an object, which is a different lookup.

Quiz

Why the return value has parentheses

A function you are reading returns (user, created) and the caller writes user, created = fetch(). What is that return value, and why is it not a list?

  1. AA list, since parentheses and square brackets are interchangeable syntax for the same sequence type
  2. BA tuple, which is immutable and hashable, and is the idiomatic shape for a fixed-length heterogeneous record
  3. CA generator, because the parentheses around a comma-separated series always create a lazy sequence
  4. DAn unpacking expression, which is a distinct type that only exists at the moment of assignment
Show answer

Correct answer: B — A tuple, which is immutable and hashable, and is the idiomatic shape for a fixed-length heterogeneous record

It is a tuple. The distinction lists-are-mutable-and-homogeneous versus tuples-are-immutable-and-positional is real and load-bearing: only the tuple can be a dict key or a set member, because only it is hashable. The generator option is the sharpest distractor because parentheses around a comprehension genuinely do produce a generator, so the rule feels like it should extend to a plain series. It does not. The comma is what makes a tuple; the parentheses are usually optional.

Quiz

Reading a comprehension

Which comprehension is the equivalent of items.filter(i => i.active).map(i => i.name)?

  1. A[i for i in items if i.active map i.name]
  2. B[for i in items if i.active: i.name]
  3. C[i.name for i in items if i.active]
  4. D[i.active and i.name for i in items]
Show answer

Correct answer: C — [i.name for i in items if i.active]

The output expression comes first, then the loop, then the filter. Reading order and execution order are deliberately different: you see what you are collecting before you see where it comes from. The and-expression option is the interesting wrong answer because it runs without error and returns a list of the same length as items, quietly containing False for every inactive item, so it is the shape of bug that reaches production.

Recall

Iterating a dict

You will read this in every config-handling function you meet, and getting it wrong gives you a silent surprise rather than an error.

What do you get when you iterate a dict directly with for x in d:, and what are the two methods that give you the other two views?

Reveal answer

Iterating a dict directly yields its KEYS, not its values and not entries. That is the surprise, because the analogous JavaScript idiom people reach for is Object.entries. The two other views are d.values() for values and d.items() for key-value tuples, which is what you almost always want: for key, value in d.items():. All three are lazy views over the live dict, not copies. And since Python 3.7 dicts preserve insertion order as a language guarantee, so iteration order is the order things were added.

Quiz

Slicing past the end

rows is a list of three items and you evaluate rows[0:10]. What happens?

  1. AYou get an empty list, because the slice is not fully satisfiable
  2. BYou get a IndexError, the same error a direct index past the end raises
  3. CYou get three items followed by seven None values, padding the slice to the requested length
  4. DYou get all three items, because slice bounds are clamped rather than checked
Show answer

Correct answer: D — You get all three items, because slice bounds are clamped rather than checked

Slices clamp silently; direct indexing does not. rows[10] raises IndexError and rows[0:10] cheerfully returns whatever is there. That asymmetry is worth holding, because it means a slice can never tell you your assumption about length was wrong. Negative indices are also real and useful: rows[-1] is the last item and rows[:-1] is everything except the last.

Sign in to track your progress →