LeapQuant proposes efficient linear attention with accurate recurrent state quantization
A new method reduces inference bottlenecks in linear‑attention models by applying precise quantization to recurrent states.
A new method reduces inference bottlenecks in linear‑attention models by applying precise quantization to recurrent states.