promptdojo_

Decoding, KV cache, and quantization — step 2 of 7

Chapter 23 showed cached prompt-prefix tokens billed at a fraction of fresh ones. Which serving mechanism from this lesson is behind that pricing?