EduVid-LLM: A Planning-Centric Architecture for Automated Educational Video Generation
Description
This preprint introduces EduVid-LLM, a novel framework for automated educational video generation using a planning-centric architecture. It employs a 'Director' LLM to create structured JSON video plans and an 'Animator' LLM to synthesize coherent Manim animations with an automated debugging loop. Designed for educational narratives, it addresses the limitations of monolithic text-to-video systems. Demo available at https://huggingface.co/spaces/gokul00060/EduVid-LLM. Code to be released on GitHub post-preprint. Aimed at enhancing accessibility for educators and content creators.
Files
main_organized.pdf
Files
(311.9 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:16ec7efed43082b2c0c1a4739c85d287
|
311.9 kB | Preview Download |
Additional details
Related works
- Is supplement to
- https://huggingface.co/spaces/gokul00060/EduVid-LLM (URL)
Dates
- Created
-
2025-08-20
Software
- Repository URL
- https://huggingface.co/spaces/gokul00060/EduVid-LLM
- Programming language
- Python
- Development Status
- Active
References
- [1] J. D. Bransford, A. L. Brown, and R. R. Cocking, How People Learn: Brain, Mind, Experience, and School. Washington, D.C.: National Academy Press, 2000. [2] A. Celikyilmaz, E. Clark, and J. Gao, "Evaluation of Text Generation: A Survey," arXiv preprint arXiv:2006.14799, 2020. [3] J. Es, R. E. de Vries, J. van der Velde, and M. de Rijke, "RAG vs Fine- tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture," arXiv preprint arXiv:2401.08406, 2023. [4] P. Lewis et al., "Retrieval-Augmented Generation for Knowledge- Intensive NLP Tasks," in Advances in Neural Information Processing Systems 33 (NeurIPS 2020). [5] P. Liang et al., "Holistic Evaluation of Language Models," arXiv preprint arXiv:2211.09110, 2022. [6] H. Luo, Z. Chen, et al., "Video-LLaMA: An Instruction-tuned Audio- Visual Language Model for Video Understanding," arXiv preprint arXiv:2306.02858, 2023. [7] 3Blue1Brown, "Manim Community," 2020. [Online]. Available: https: //www.manim.community/. [8] R. E. Mayer, Multimedia Learning (2nd ed.). Cambridge University Press, 2009. [9] J. S. Park et al., "Generative Agents: Interactive Simulacra of Human Behavior," in Proc. 36th Annual ACM Symposium on User Interface Software and Technology (UIST '23), New York, NY, USA: ACM, 2023, Article 2, pp. 1–22. [10] U. Singer et al., "Make-A-Video: Text-to-Video Generation without Text-Video Data," in International Conference on Learning Representa- tions (ICLR 2023). [11] L. Wang, C. Ma, X. Feng, et al., "A Survey on Large Language Model based Autonomous Agents," arXiv preprint arXiv:2308.11432, 2023. [12] T. Wu et al., "VideoDirectorGPT: Consistent Multi-scene Video Gener- ation via LLM-Guided Planning," arXiv:2309.15091, 2023. [13] L. Yu et al., "VideoPoet: A Large Language Model for Zero-Shot Video Generation," Google Research, 2023. [14] Q. Wu et al., "AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation," arXiv:2308.08155, 2023.