Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answer
Description
Recent advances in multimodal large language models (LLMs) have shown extreme effectiveness in visual question answering (VQA). However, the design nature of these end-to-end models prevents them from being interpretable to humans, undermining trust and applicability in critical domains. While post-hoc rationales offer certain insight into understanding model behavior, these explanations are not guaranteed to be faithful to the model. In this paper, we address these shortcomings by introducing an interpretable by design model that factors model decisions into intermediate human-legible explana
Research goal: Does SMoES-style dynamic modality routing improve out-of-distribution robustness on VCR adversarial splits compared to Top-2 modality-agnostic MoE-VLMs with matched expert counts?
Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 7.8/10.
Notes
Files
paper.pdf
Files
(84.4 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:7afb60aa17f8470b0ef548baa7d5e8e1
|
84.4 kB | Preview Download |