vLLM
When Memory Layout Assumptions Collide: Fixing a KV Cache Bug in vLLM
Two independently correct components — KV offloading and per-token-head quantization — corrupted generations the moment they were combined. The fix was four lines. Finding where those four lines belonged took a trip through GPU memory layouts.




