Published August 3, 2023 | Version v1

Towards a performance analysis on pre-trained Visual Question Answering models for autonomous driving

  • 1. University of Limerick
  • 2. Valeo Vision Systems

Description

This short paper presents a preliminary analysis of three popular Visual Question Answering (VQA) models, namely ViLBERT, ViLT, and LXMERT, in the context of answering questions relating to driving scenarios. The performance of these models is evaluated by comparing the similarity of responses to reference answers provided by computer vision experts. Model selection is predicated on the analysis of transformer utilization in multimodal architectures.  The results indicate that models incorporating cross-modal attention and late fusion techniques exhibit promising potential for generating improved answers within a driving perspective. This initial analysis serves as a launchpad for a forthcoming comprehensive comparative study involving nine VQA models and  sets the scene for further investigations into the effectiveness of VQA model queries in self-driving scenarios. Supplementary material is available on the Github page. 
 

Files

poster.pdf

Files (2.0 MB)

Name Size Download all
md5:793f84bd7c69416e440054d134462bd8
1.1 MB Preview Download
md5:1ff251b113c6d8009cb08c262a9aca7b
463.8 kB Preview Download
md5:004e178b8ce0d5241512f53ec9542239
503.7 kB Preview Download

Additional details

Funding

Science Foundation Ireland
SFI Centre for Research Training in Foundations of Data Science 18/CRT/6049