How LLMs Work
Book Labs
All labs
21 Book labs · Build it by hand

Backprop, one step at a time

From Chapter 21 of Large Language Models from the Ground Up. When a value is used in more than one place, its gradient is the sum of what comes back along every route. Step the sweep and watch the two routes arrive one at a time.

Interactive

One input, two routes, one loss

a feeds both b = a × 2 and c = a + 5, and they meet again at L = b × c. Press Step to run one _backward() call at a time.

0
a.grad so far. The true rate is 22, and the sweep only gets there once both routes have paid in.not started
What to notice

Neither route is the answer on its own. Route 1 delivers 8 × 2 and route 2 delivers 6 × 1, and only their sum matches what you get by nudging a and re-running the arithmetic. That is why every gradient starts at 0 and every rule in the engine uses += rather than =: had the second route assigned instead of added, it would have erased the first, and the answer would be 6 instead of 22 — a number with nothing obviously wrong about it. It is also why the sweep has to run in reverse topological order: a node may only hand its gradient onward once every operation that uses it has already contributed.

← Back to all labs