MS-GLA: Multi-Scale Gated Linear Attention for Transformers
A new pre-print introduces Multi-Scale Gated Linear Attention to address representational bottlenecks across multiple temporal resolutions in transformers.
A new pre-print introduces Multi-Scale Gated Linear Attention to address representational bottlenecks across multiple temporal resolutions in transformers.