World Wires · story 8595 · corroborated · 1 source(s)

Mira: Memory‑Efficient MoE Inference Using Adaptive Caching and Predictive Expert Staging

Mira introduces adaptive caching and predictive expert staging to reduce memory consumption during inference of mixture‑of‑experts models on single‑GPU sys

Open in the desk

Coverage

What this site indexes