Disable cross-layer KV blocks for per-token-head quant
Fixed silent KV-cache corruption when OffloadingConnector was combined with a per-token-head quantized KV cache.
Per-token-head quant stores an inline fp32 scale in each cell’s padded tail, and the attention backend carves scale views assuming a per-layer contiguous buffer — but the offload path allocates one cross-layer interleaved buffer shared by all layers, so the scales aliased neighbouring K/V data from the first decode. After review the guard moved out of the connector into use_uniform_kv_cache(), covering every connector in four lines. Full write-up on the blog.