Published June 2, 2026 | Version v1

VISION-LANGUAGE MODELS IN THE ERA OF MULTIMODAL FOUNDATION MODELS

  • 1. Tashkent University of Information Technologies named after Muhammad al-Khwarizmi, Tashkent, Uzbekistan

Description

Vision-language models (VLMs) have moved from task-specific image-text encoders to general multimodal foundation models capable of visual reasoning, captioning, retrieval, question answering, grounding, optical character recognition, instruction following and multi-image interaction. This review analyzes twenty influential papers and technical reports published mainly between 2021 and 2025, including CLIP, ALIGN, ALBEF, BLIP, Flamingo, CoCa, PaLI, BLIP-2, LLaVA, MiniGPT-4, InstructBLIP, Kosmos-2, Qwen-VL, CogVLM, GPT-4V, Gemini, InternVL, MM1, PaliGemma and Molmo/PixMo. The purpose is to identify how the field has changed in architecture, data construction, training objectives, evaluation practice and deployment challenges. The review shows that early contrastive alignment created transferable visual representations, while recent models increasingly connect strong vision encoders with large language models through lightweight adapters, query transformers, visual experts, instruction tuning and interleaved multimodal data. Current progress is driven not only by model scale, but also by the quality of captions, grounding data, OCR-rich samples, instruction datasets and safety evaluation. The main unresolved problems remain visual hallucination, weak spatial grounding, limited transparency of training data, high computational cost, benchmark saturation, and insufficient reliability in high-stakes domains. The paper concludes that the next stage of VLM research will depend on data-centric training, verifiable grounding, efficient open models, and evaluation protocols that measure real-world visual reasoning rather than only benchmark accuracy.

Files

22-30.pdf

Files (225.8 kB)

Name Size Download all
md5:293c34c8e9f6fa63a9ec145b646be6b2
225.8 kB Preview Download

Additional details

References

  • [1] Radford, A., Kim, J. W., Hallacy, C., et al. Learning transferable visual models from natural language supervision. Proceedings of the 38th International Conference on Machine Learning, PMLR 139, 8748–8763, 2021.
  • [2] Jia, C., Yang, Y., Xia, Y., et al. Scaling up visual and vision-language representation learning with noisy text supervision. Proceedings of the 38th International Conference on Machine Learning, PMLR 139, 4904–4916, 2021.
  • [3] Li, J., Selvaraju, R. R., Gotmare, A. D., Joty, S., Xiong, C., and Hoi, S. Align before fuse: Vision and language representation learning with momentum distillation. Advances in Neural Information Processing Systems, 34, 9694–9705, 2021.
  • [4] Li, J., Li, D., Xiong, C., and Hoi, S. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. Proceedings of the 39th International Conference on Machine Learning, 2022.
  • [5] Alayrac, J. B., Donahue, J., Luc, P., et al. Flamingo: A visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35, 23716–23736, 2022.