Published January 1, 2025
| Version v1
Conference paper
Open
Towards Efficient Keyframe Selection for News Video Captioning
Authors/Creators
Description
Vision-language models are increasingly used for video captioning, but their reliance on uniform frame sampling introduces inefficiencies. Uniform sampling risks overlooking informative frames rich in textual or structural cues while redundantly processing visually similar frames, leading to inflated computational costs and degraded caption quality. To address this, we propose a layout-aware keyframe selection method that integrates shot segmentation with layout detection to identify frames containing the most semantically informative visual and textual elements. By prioritizing event-rich frames, our approach reduces redundancy while preserving critical contextual information for downstream captioning. On a benchmark of English-language news videos, our method achieved an event matching F1 score of 0.961, a 14\
Files
2025-Towards_Efficient_Keyframe_Selection_for_News_Video_Captioning.pdf
Files
(2.4 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:8699694d722c5e54b0bf0d7f4d5f024e
|
2.4 MB | Preview Download |