TurboQuant and the LLM KV cache VRAM bottleneck
Google’s TurboQuant compresses the LLM KV cache 5x with near-zero accuracy loss. How PolarQuant’s randomized rotation flattens outliers, and how QJL residual correction keeps the attention dot product unbiased.