Time travel, and the Ark mapping
Time travel means reading and forking history.
const snapshot = await graph.getState(config); // latest state + next node
for await (const s of graph.getStateHistory(config)) {} // walk every checkpoint
// re-enter at a specific past checkpoint:
const past = { configurable: { thread_id: "…", checkpoint_id: "…" } };
await graph.getState(past);
// fork: write new values onto a checkpoint, creating a new branch:
await graph.updateState(config, { retryCount: 0 });getState returns the state and what node runs next. getStateHistory is the full lineage. updateState forks: it writes a new checkpoint with your override, which is how you replay from a modified past.
The build’s written mapping is the deliverable: session.json maps to thread, writeback maps to checkpoint, incident-log maps to history. Once you have built it, you can say whether Ark’s bespoke files are earning their keep or reinventing the checkpointer.
Try it yourself
What forking actually means
updateState is described as "forking". What actually forks, and how does that differ from getState with a checkpoint_id?
Show answer
Correct answer: C — updateState writes a new checkpoint on top of an existing one (forking history); getState with checkpoint_id only reads
getState is read-only time travel: point it at a checkpoint_id and it shows you that moment. updateState is a write: it creates a new checkpoint carrying your overridden values, branching the history from that point forward -- that's the fork.
The honest equivalent to Ark's incident log
The mapping is close enough to be useful and wrong enough to be worth naming. Both halves are the answer.
What is the LangGraph equivalent of Ark's incident-log.jsonl, and what's the limit of that analogy?
Reveal answer
getStateHistory / the checkpoint lineage is the equivalent: an ordered record you can walk to reconstruct what happened. The analogy breaks down because Ark's incident log is a deliberately curated, human-readable forensic record with editorial judgment about what's worth logging, while checkpoint history is an automatic, complete record of every state transition -- more like a raw audit trail than a curated narrative.
Add a Postgres checkpointer, then kill and resume mid-run
Put persistence under the module 3 impulse graph and prove it survives a crash.
- Extend ~/dev/groundschool/builds/m3-impulse-graph with a PostgresSaver checkpointer.
- Compile the graph with { checkpointer } and run checkpointer.setup() once.
- Start a run with a fixed thread_id, then kill the process partway through (e.g. mid-modulate).
- Start a new process with the same thread_id and config, and invoke with null inputs to resume.
- Confirm the run completes as a resume, not a fresh run from START.
- Walk graph.getStateHistory(config) and annotate: which checkpoint is the writeback equivalent, which is the boot equivalent.
- Write the mapping: session.json -> thread, writeback -> checkpoint, incident-log -> history.
Mid-run kill and resume verified
Confirm the resume behavior after a mid-run process kill.
Killing the process mid-graph and resuming with the same thread_id completes the same thread (not a fresh run), and the checkpoint lineage is annotated with which checkpoint corresponds to a writeback and which to a boot.