Retention and Transmission of Information in Sliding-Window KV Inference
Sliding-window KV inference processes sequences incrementally with a fixed-size cache of recent key and value states without additional training.
Sliding-window KV inference processes sequences incrementally with a fixed-size cache of recent key and value states without additional training.