Published January 1, 2025 | Version v1

Towards Efficient Keyframe Selection for News Video Captioning

Description

Vision-language models are increasingly used for video captioning, but their reliance on uniform frame sampling introduces inefficiencies. Uniform sampling risks overlooking informative frames rich in textual or structural cues while redundantly processing visually similar frames, leading to inflated computational costs and degraded caption quality. To address this, we propose a layout-aware keyframe selection method that integrates shot segmentation with layout detection to identify frames containing the most semantically informative visual and textual elements. By prioritizing event-rich frames, our approach reduces redundancy while preserving critical contextual information for downstream captioning. On a benchmark of English-language news videos, our method achieved an event matching F1 score of 0.961, a 14\

Files

2025-Towards_Efficient_Keyframe_Selection_for_News_Video_Captioning.pdf

Additional details