SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs
Description
Mixture-of-Experts (MoE) has become a prevalent backbone for large vision-language models (VLMs), yet how modality-specific signals should guide expert routing remains under-explored. Existing routing strategies are either hand-crafted or modality-agnostic, relying on idealized priors that ignore the layer-dependent modality fusion patterns in MoE-VLMs and provide little guidance for expert specialization. We propose Soft Modality-guided Expert Specialization (SMoES), which consists of dynamic soft modality scores that capture layer-dependent fusion patterns, an expert binning mechanism aligne
Research goal: What is the impact of varying expert count in SMoES-based MoE-VLMs on MMMU score variance and expert routing robustness compared to dense VLMs?
Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 8.2/10.
Notes
Files
paper.pdf
Files
(81.4 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:2ef2ec1cbe484f17c307e0eaf45fc736
|
81.4 kB | Preview Download |