Author

Date of Award

11-2025

Document Type

Dissertation

Publisher

Santa Clara : Santa Clara University, 2025

Degree Name

Master of Science (MS)

Department

Computer Science and Engineering

First Advisor

Oana Ignat

Abstract

With the rapid development of text-to-video (T2V) generation models, the evaluation and benchmarking focus has shifted from merely visual fidelity to semantic faithfulness, among which one key dimension is the cultural understanding and representation abilities. However, current T2V benchmarks focus primarily on single culture scenes, lacking systematic study of cross-cultural (i.e., multiple cultures within one T2V prompt) generation. Besides, there lack multi-agent frameworks for improving cultural fidelity of T2V generation.

To fill this gap, we propose CRAFT (MultiCultuRal Multi-Agent Framework for Textto- video Generation), a multi-agent prompt refinement framework for improving cultural fidelity. We design 3 agent architectures: single-agent (SA) uses a single agent to refine the whole prompt, multi-agent sequential (MAS), and multi-agent parallel (MAP) use specialized agents taking charge of person, action, and location dimension refinement respectively. We construct a dataset of 972 videos, consisting of 3 cultures (Chinese, American, Romanian), 3 cultural action types (food, music, dance), totaling 27 cultural actions, and 9 iconic cultural locations, covering both mono-cultural and cross-cultural scenarios. Evaluation framework that combines both CLIP-based automatic metrics and VLM-based judgement, evaluates the generated videos from multiple dimensions spanning cultural relevance, visual similarity, and text-image alignment.

Experimental results show that: 1) MAP achieves the highest cultural relevance among all refinement pipelines, and at the same time achieves the most balanced improvement on both video quality and temporal consistency; 2) cross-cultural prompts pose a larger challenge than mono-cultural prompts for T2V models, as illustrated by lower cultural relevance; 3) CLIP-based metrics and VLM-based evaluation show moderate to strong positive correlation on evaluating cultural relevance.

Our research provides a feasible framework for building culturally faithful and inclusive T2V generation systems, paving ways for more respectful and multi-cultural video content generation in the context of globalization. Our dataset is at https://huggingface.co/datasets/sl-scu/multicultural_ multiagent_videos, also complete code and dataset is available at https://github. com/AIM-SCU/CRAFT.

Available for download on Thursday, September 28, 2028

Share

COinS