Looped MoE: Flattening Experts and Untying Attention
A study explores combining looped transformers and mixture-of-experts architectures to improve expert utilization and parameter efficiency.
A study explores combining looped transformers and mixture-of-experts architectures to improve expert utilization and parameter efficiency.