lesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blockslesson 3 of 5 · attention and transformer blocks
Doubling a transformer's context window from 100k to 200k tokens does what to attention's compute cost, and why?