Date of Award

6-9-2026

Document Type

Thesis

Publisher

Santa Clara : Santa Clara University, 2026

Department

Electrical and Computer Engineering

First Advisor

Tokunbo Ogunfunmi

Abstract

Over the past decade, advances in embedded computing, robotics frameworks, and AI have significantly lowered the barrier to developing intelligent robotic systems. Some single board computers such as the Raspberry Pi are now capable of running operating systems like Ubuntu and frameworks such as ROS2, allowing us to build robots that can perform tasks including autonomous navigation, computer vision, sensor fusion, and speech interaction. As a result, technologies that were once limited to large research labs are becoming increasingly accessible in educational and lower level research environments.

At the same time, recent developments in AI have changed expectations for how robots interact with humans. Traditional robotic systems often rely on predefined buttons, remote controllers, or rigid command structures. In contrast, modern intelligent systems are expected to understand natural language, interpret user intent, and respond in a more flexible and intuitive manner. This trend is particularly important in education, where students often learn topics such as machine learning, computer vision, and robotics separately without seeing how these technologies can be integrated into a complete physical system.

Although many educational robotic platforms are available today, most low-cost systems focus on basic functions such as line following or obstacle avoidance. More advanced platforms frequently depend on proprietary software or cloud services that provide limited transparency for learning. As a result, students may gain experience with individual components but have fewer opportunities to understand how perception, decision-making, and control work together in a real robotic application.

Another challenge is that many voice-controlled systems are primarily designed for standard languages and accents. In multilingual regions such as China, a large number of people use regional dialects in their daily communication. While these dialects remain important parts of local culture and identity, they are often underrepresented in voice-based AI applications. This creates an opportunity to explore how modern speech recognition and language understanding technologies can improve accessibility and inclusiveness in human-robot interaction. Supporting dialect-based interaction can help make intelligent systems more natural and approachable for a wider range of users.

To address these challenges, we developed a Chinese dialect-controlled multifunctional robot built on a Raspberry Pi 5 and ROS2 architecture. The system combines computer vision (CV), ultrasonic sensing, speech recognition, large language model (LLM) understanding, and robot motion control into a single platform. In addition to practical robotic functions such as color detection, visual line following, and obstacle avoidance, the robot is capable of understanding voice commands spoken in Cantonese and Sichuanese dialects and converting them into executable actions. Rather than focusing solely on performance optimization, the project aims to provide an educational example of how modern AI technologies can be integrated into a complete robotic system. Furthermore, the platform is designed to be modular and extensible so that future students can continue improving and expanding its capabilities.

Share

COinS