ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching
Description
Sparse Mixture-of-Experts (MoE) models can outperform dense large language models at similar computation by activating only a small set of experts per token. However, stacking many expert modules introduces substantial parameter memory, which makes MoE models difficult to deploy in memory-constrained environments such as single-GPU devices. Offloading alleviates this issue by storing inactive experts in CPU memory and loading them on demand, but existing methods remain limited: static caches disregard input-dependent routing, and methods that train separate models to predict expert usage ahead
Research goal: What is the inference throughput (tokens per second) and FLOPs efficiency of SMoES-based multimodal models relative to dense baselines when evaluated on cross-modal distribution shifts in the MMMU benchmark?
Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 7.5/10.
Notes
Files
paper.pdf
Files
(85.9 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:51180015360d17635373c295ddf2988e
|
85.9 kB | Preview Download |