Date of Award
6-2026
Document Type
Thesis
Publisher
Santa Clara : Santa Clara University, 2026
Degree Name
Master of Science (MS)
Department
Electrical and Computer Engineering
First Advisor
Hoeseok Yang
Abstract
Efficient inference in Large Language Models (LLMs) requires deciding where computation can be reduced while preserving model quality. We study this problem through multilayer perceptron (MLP) activation modification, activation sparsification and token-level conditional routing. We propose Sensitivity-Aware Thresholding for Sparsity (SATS), a threshold calibration method to choose layerwise gate thresholds using a local MLP output sensitivity proxy rather than calibrating thresholds directly from activation percentiles. While SATS retains the existing mechanism of sparsifying MLP activations by thresholding gate activations, it replaces percentile-based calibration with a sensitivity-aware selection rule. We then introduce a lightweight token routing framework that dynamically selects between a base path and a modified path on a per-token basis, rather than applying the modified computation uniformly to all tokens. We evaluate both methods on multiple recent open-weight LLMs. We also study activation modification methods covering full and partial ReLU-fication, with both full weight and LoRA-based recovery. We analyzed layerwise sensitivity and demonstrated the motivation behind threshold-based activation modification as our main design path. Our results show that threshold-based activation modification offers a more favorable accuracy-sparsity tradeoff than static ReLU-fication. The results also demonstrate that our proposed SATS improves over the threshold-based sparsification baseline at matched actual sparsity and that token routing yields a more favorable quality-throughput trade-off than static activation modification baselines. Overall, our results suggest that activation modification, improved threshold calibration and token routing can improve the quality-throughput trade-off in LLMs.
Recommended Citation
Paul, Bishmoy, "Sensitivity-Aware Thresholding and Token Routing for Activation Sparsification in Large Language Models" (2026). Electrical and Computer Engineering Master's Theses. 16.
https://scholarcommons.scu.edu/elec_mstr/16
