Published August 20, 2025 | Version 1.0

EduVid-LLM: A Planning-Centric Architecture for Automated Educational Video Generation

  • 1. ROR icon Sathyabama Institute of Science and Technology

Description

This preprint introduces EduVid-LLM, a novel framework for automated educational video generation using a planning-centric architecture. It employs a 'Director' LLM to create structured JSON video plans and an 'Animator' LLM to synthesize coherent Manim animations with an automated debugging loop. Designed for educational narratives, it addresses the limitations of monolithic text-to-video systems. Demo available at https://huggingface.co/spaces/gokul00060/EduVid-LLM. Code to be released on GitHub post-preprint. Aimed at enhancing accessibility for educators and content creators.

Files

main_organized.pdf

Files (311.9 kB)

Name Size Download all
md5:16ec7efed43082b2c0c1a4739c85d287
311.9 kB Preview Download

Additional details

Related works

Dates

Created
2025-08-20

Software

Repository URL
https://huggingface.co/spaces/gokul00060/EduVid-LLM
Programming language
Python
Development Status
Active

References

  • [1] J. D. Bransford, A. L. Brown, and R. R. Cocking, How People Learn: Brain, Mind, Experience, and School. Washington, D.C.: National Academy Press, 2000. [2] A. Celikyilmaz, E. Clark, and J. Gao, "Evaluation of Text Generation: A Survey," arXiv preprint arXiv:2006.14799, 2020. [3] J. Es, R. E. de Vries, J. van der Velde, and M. de Rijke, "RAG vs Fine- tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture," arXiv preprint arXiv:2401.08406, 2023. [4] P. Lewis et al., "Retrieval-Augmented Generation for Knowledge- Intensive NLP Tasks," in Advances in Neural Information Processing Systems 33 (NeurIPS 2020). [5] P. Liang et al., "Holistic Evaluation of Language Models," arXiv preprint arXiv:2211.09110, 2022. [6] H. Luo, Z. Chen, et al., "Video-LLaMA: An Instruction-tuned Audio- Visual Language Model for Video Understanding," arXiv preprint arXiv:2306.02858, 2023. [7] 3Blue1Brown, "Manim Community," 2020. [Online]. Available: https: //www.manim.community/. [8] R. E. Mayer, Multimedia Learning (2nd ed.). Cambridge University Press, 2009. [9] J. S. Park et al., "Generative Agents: Interactive Simulacra of Human Behavior," in Proc. 36th Annual ACM Symposium on User Interface Software and Technology (UIST '23), New York, NY, USA: ACM, 2023, Article 2, pp. 1–22. [10] U. Singer et al., "Make-A-Video: Text-to-Video Generation without Text-Video Data," in International Conference on Learning Representa- tions (ICLR 2023). [11] L. Wang, C. Ma, X. Feng, et al., "A Survey on Large Language Model based Autonomous Agents," arXiv preprint arXiv:2308.11432, 2023. [12] T. Wu et al., "VideoDirectorGPT: Consistent Multi-scene Video Gener- ation via LLM-Guided Planning," arXiv:2309.15091, 2023. [13] L. Yu et al., "VideoPoet: A Large Language Model for Zero-Shot Video Generation," Google Research, 2023. [14] Q. Wu et al., "AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation," arXiv:2308.08155, 2023.