Date of Award
11-2025
Document Type
Dissertation
Publisher
Santa Clara : Santa Clara University, 2025
Degree Name
Master of Science (MS)
Department
Computer Science and Engineering
First Advisor
Oana Ignat
Abstract
With the rapid development of text-to-video (T2V) generation models, the evaluation and benchmarking focus has shifted from merely visual fidelity to semantic faithfulness, among which one key dimension is the cultural understanding and representation abilities. However, current T2V benchmarks focus primarily on single culture scenes, lacking systematic study of cross-cultural (i.e., multiple cultures within one T2V prompt) generation. Besides, there lack multi-agent frameworks for improving cultural fidelity of T2V generation.
To fill this gap, we propose CRAFT (MultiCultuRal Multi-Agent Framework for Textto- video Generation), a multi-agent prompt refinement framework for improving cultural fidelity. We design 3 agent architectures: single-agent (SA) uses a single agent to refine the whole prompt, multi-agent sequential (MAS), and multi-agent parallel (MAP) use specialized agents taking charge of person, action, and location dimension refinement respectively. We construct a dataset of 972 videos, consisting of 3 cultures (Chinese, American, Romanian), 3 cultural action types (food, music, dance), totaling 27 cultural actions, and 9 iconic cultural locations, covering both mono-cultural and cross-cultural scenarios. Evaluation framework that combines both CLIP-based automatic metrics and VLM-based judgement, evaluates the generated videos from multiple dimensions spanning cultural relevance, visual similarity, and text-image alignment.
Experimental results show that: 1) MAP achieves the highest cultural relevance among all refinement pipelines, and at the same time achieves the most balanced improvement on both video quality and temporal consistency; 2) cross-cultural prompts pose a larger challenge than mono-cultural prompts for T2V models, as illustrated by lower cultural relevance; 3) CLIP-based metrics and VLM-based evaluation show moderate to strong positive correlation on evaluating cultural relevance.
Our research provides a feasible framework for building culturally faithful and inclusive T2V generation systems, paving ways for more respectful and multi-cultural video content generation in the context of globalization. Our dataset is at https://huggingface.co/datasets/sl-scu/multicultural_ multiagent_videos, also complete code and dataset is available at https://github. com/AIM-SCU/CRAFT.
Recommended Citation
Li, Shuowei, "CRAFT: A Multi-Agent Framework for Multicultural Text-to-Video Generation" (2025). Computer Science and Engineering Master's Theses. 65.
https://scholarcommons.scu.edu/cseng_mstr/65
