Mira: Memory‑Efficient MoE Inference Using Adaptive Caching and Predictive Expert Staging
Mira introduces adaptive caching and predictive expert staging to reduce memory consumption during inference of mixture‑of‑experts models on single‑GPU sys
Mira introduces adaptive caching and predictive expert staging to reduce memory consumption during inference of mixture‑of‑experts models on single‑GPU sys