lesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blocks
Checkpoint
One last thing before we move on. pass this to mark the lesson done, or skip and keep moving. hop to the next when you're ready.
Add the causal mask — the decoder rule that keeps next-token training from cheating. For each token i of 3, compute scaled dot-product scores against ONLY keys j <= i, softmax over those, and pad the row with 0.0 up to length 3. Print f"token {i}: {row}" (weights rounded to 2) and finish with the literal line "upper triangle is zero: no token sees the future".