Author

Date of Award

6-2026

Document Type

Dissertation

Publisher

Santa Clara : Santa Clara University, 2026

Degree Name

Doctor of Philosophy (PhD)

Department

Computer Science and Engineering

First Advisor

Xiang Li

Abstract

Power grid cascading failures, sequences of overload-triggered line trips that propagate across a transmission network, are among the most consequential failure modes in modern infrastructure. Identifying the transmission lines whose failure is most likely to initiate or sustain a large cascade is a vulnerability-analysis problem whose combinatorial intractability has driven prior work toward heuristic contingency screening, stochastic simulation, and classical machine learning. None provides a unified learning framework that scales across grids while producing interpretable vulnerability rankings.

This dissertation develops three contributions, each establishing a successively stronger property of attention extracted from a model trained on cascade data.

The first contribution (Chapter 3) establishes that power-grid cascading failure can be cast as a sequence-modeling problem learnable by a Transformer, and that attention extracted from the trained model carries meaningful structure usable for downstream analysis. An encoder-decoder Transformer trained on cascade trajectories recovers the simulator’s vulnerability rankings from initial-failure tokens alone, and an Independent Cascade approximation derived from its attention matrix runs orders of magnitude faster than power-flow simulation on the 852-line SciGrid benchmark.

Building on Chapter 3, the second contribution (Chapter 4) shows that extracted attention is itself the vulnerability signal. An encoder-only Transformer trained self-supervised on the Dual Representation of cascade state yields Initiatives and Passives, read directly from the masked self-attention matrix, which outperform two physics-informed baselines across four evaluation scenarios on three European national grids from PyPSA-EUR.

The third contribution (Chapter 5) shows that extracted attention transfers across grids, lifting the per-grid retraining limitation of Chapters 3 and 4. CG-CAE, a GRU-gated graph attention network trained once on three grids, transfers zero-shot to six structurally unseen grids, where its attention coefficients aggregate into a cascade exposure score per line. Mean percentile rank of high-exposure lines improves by roughly 14 percentage points over the better physics-informed baseline on every evaluation grid.

Together, the three contributions establish that attention extracted from cascade-trained models is a meaningful, vulnerability-bearing, and cross-grid transferable signal for power-grid vulnerability analysis.

Available for download on Thursday, September 28, 2028

Share

COinS