Date of Award
6-2026
Document Type
Dissertation
Publisher
Santa Clara : Santa Clara University, 2026
Degree Name
Doctor of Philosophy (PhD)
Department
Computer Science and Engineering
First Advisor
Yi Fang
Abstract
Artificial intelligence and machine learning systems are deeply embedded in modern information ecosystems, yet the bias they exhibit is not a single phenomenon. This dissertation identifies three distinct stages of the AI pipeline at which bias surfaces, each demanding a di!erent conceptual treatment: measurement bias in preprocessing, addressed through identification fairness; content bias at the corpus level, addressed through representation equity; and systemic bias in retrieval, addressed through distributional justice. Measurement bias arises when raw signals about people are converted into categorical labels, and errors at this stage propagate into every subsequent computation. In race and ethnicity prediction from names, models trained on imbalanced data identify majority groups accurately while systematically misidentifying minorities. This dissertation introduces Fairness-aware Race and Ethnicity Detection (FRED), the first framework to embed fairness constraints directly into the training of name-based demographic inference models, showing that the group-level performance gap can be substantially closed without sacrificing accuracy. Content bias arises when the distribution of who is quoted or covered diverges from equitable representation. Existing analyses of news source diversity are retrospective, labor-intensive, and limited to a single demographic dimension. This dissertation presents DIANES: A Diversity, Equity, and Inclusion Audit Toolkit for News Sources, the first end-to-end system to unify quote extraction, speaker identification, and joint gender and race/ethnicity inference in a near-real-time pipeline delivered through a WordPress plugin, dashboards, and APIs.
Systemic bias arises in the retrieval layer of retrieval-augmented AI agents, where the retriever may under-expose long-tail knowledge items before any answer is produced. No prior work has extended classical retrievability theory to agent settings where queries are generated dynamically during multi-step reasoning. This dissertation develops, for the first time, a formal framework for measuring retrievability in AI agents, introducing utilization-aware metrics and group-level disparity measures over agent access traces. An empirical study on HotpotQA across 7,405 tasks and 66,581 knowledge items reveals that even a near-perfect agent reaches only about 43% of the corpus at K = 3, leaving the majority e!ectively invisible to aggregate accuracy metrics. Together, these three contributions provide both the language and the tools to pursue fairness at the modeling, system, and retrieval layers of the AI pipeline simultaneously.
Recommended Citation
Shang, Xiaoxiao, "Fairness, Representation, and Retrievability in AI Information Systems" (2026). Engineering Ph.D. Theses. 68.
https://scholarcommons.scu.edu/eng_phd_theses/68
