Published August 26, 2026 | Version v1

Cached Descriptive Asset Files (CDAF): A Sidecar Format for Token-Efficient Video Understanding in Agentic Pipelines

  • 1. ROR icon Indian Institute of Technology Patna

Description

Multimodal language models can answer questions about video, but agentic workflows often pay the full video-processing cost every time the same asset is reused. We present Cached Descriptive Asset Files (CDAF), an open, model-agnostic sidecar convention that persists a timestamped description alongside a video and binds that description to the exact video bytes using SHA-256. An agent verifies freshness and reads a few hundred text tokens rather than repeatedly exposing the model to the video.

In a reproducible synthetic benchmark containing four 10–12 second videos and 20 objective questions per condition, CDAF-mediated question answering achieved 20/20 accuracy, compared with 19/20 for direct video analysis, while reducing mean prompt tokens from 3,066 to 303 (10.1×) and mean latency from 3.46 to 2.24 seconds (35.2%). Conservative accounting places the one-time generation break-even at approximately 1.3 downstream questions per video.

We specify the file format, freshness semantics, agent decision policy, and evaluation harness; analyze how CDAF converts footage selection and cut-list generation into text retrieval; and identify failure modes involving omitted detail, malicious descriptions, and synthetic-benchmark realism. The result is a deliberately simple systems primitive: video understanding becomes a reusable, verifiable text asset rather than a repeated multimodal inference cost.

Files

Cached Descriptive Asset Files PREPRINT.pdf

Files (572.6 kB)

Name Size Download all
md5:e8c2dcfdeb6dd90d59c88210fb60fabc
572.6 kB Preview Download

Additional details

Dates

Available
2026-08-26

Software

Repository URL
https://github.com/UditAkhourii/cdaf
Programming language
Python
Development Status
Active

References

  • A. Yang, A. Nagrani, P. H. Seo, A. Miech, J. Pont- Tuset, I. Laptev, J. Sivic, and C. Schmid, "Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning," in Proceedings of CVPR, 2023. [2] Gemini Team, "Gemini 1.5: Unlocking Multimodal Understanding Across Millions of Tokens of Context," arXiv:2403.05530, 2024. [3] C. Fu et al., "Video-MME: The First-Ever Compre- hensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis," arXiv:2405.21075, 2024. [4] K. Li et al., "MVBench: A Comprehensive Multi- modal Video Understanding Benchmark," in Proceed- ings of CVPR, 2024. [5] F. Bang, "GPTCache: An Open-Source Semantic Cache for LLM Applications Enabling Faster Answers and Cost Savings," in Proceedings of NLP-OSS, pp. 212–218, 2023. [6] Anthropic, "Prompt Caching," Claude Platform Doc- umentation, 2026. https://docs.anthropic.com/ en/docs/build-with-claude/prompt-caching. [7] ISO, "ISO 16684-1:2012: Graphic Technology— Extensible Metadata Platform (XMP) Specification— Part 1," 2012. [8] ISO/IEC, "ISO/IEC 15938-1:2002: Information Technology—Multimedia Content Description Interface—Part 1: Systems," 2002. [9] National Institute of Standards and Technology, "FIPS PUB 180-4: Secure Hash Standard," 2015. [10] Google, "Gemini API: Video Understanding," Google AI for Developers, updated August 17, 2026. https://ai.google.dev/gemini-api/docs/ video-understanding. [11] Remotion, "Make Videos Programmatically," 2026. https://www.remotion.dev/. 7