Date of Award
6-2-2026
Document Type
Thesis
Publisher
Santa Clara : Santa Clara University, 2026
Department
Computer Science and Engineering
First Advisor
Yuhong Liu
Abstract
The proliferation of sensitive Personally Identifiable Information (PII) on dark web marketplaces has created an urgent need for robust data protection systems, especially for vulnerable populations such as minors. Traditional PII redaction often fails to identify implicit privacy risks—such as author gender indicators or non-fictional child-related context—hidden within large-scale e-commerce datasets. This paper presents JSD, a dual-stage framework for the detection and protection of sensitive text data. The Detection phase utilizes Transformer and CNN-based architectures and Human-in-the-Loop AI to surpass the "semantic ceiling" of traditional NER approaches, enabling context-aware identification of implicit PII. The Protection phase introduces GASE (Genetic Algorithm Synthetic Evaluation), an evolutionary approach to optimizing synthetic datasets. By leveraging genetic algorithms to mutate and select for both privacy and utility, the system ensures that synthetic data remains high-fidelity for downstream machine learning tasks while eliminating critical memorization risks[3]. Experimental results demonstrate that the GASE optimization achieved a 91.71% similarity in model performance (TSTR vs. TRTR f1 ratios) compared to original datasets, while successfully reducing critical memorized PII instances to zero.
Recommended Citation
Lane, Jeffrey; Wang, Scott; Chang, Vincent; and Zhang, Bojun, "JSD: Novel Methodology for Synthetic Data Evaluation" (2026). Computer Science and Engineering Senior Theses. 364.
https://scholarcommons.scu.edu/cseng_senior/364
