WUSH‑KV: KV Cache Quantization with Data‑Adaptive Transforms
WUSH‑KV introduces a low‑bit quantization technique for transformer KV caches using data‑adaptive transforms to reduce memory and bandwidth costs.
WUSH‑KV introduces a low‑bit quantization technique for transformer KV caches using data‑adaptive transforms to reduce memory and bandwidth costs.