Author ORCID Identifier:
Date of Graduation
7-2026
Document Type
Thesis
Degree Name
Master of Science in Computer Science (MS)
Degree Level
Graduate
Department
Computer Science & Computer Engineering
Advisor/Mentor
Li, Qinghua
Committee Member
Wu, Xintao
Second Committee Member
Pan, Yanjun
Keywords
Backdoor-Sensitive Layers, Pre-trained Language Models (PLM), Secure Fine-Tuning, localizing, repairing
Abstract
Task-agnostic backdoor attacks can contaminate pre-trained language models (PLMs) in a way that survives downstream adaptation, even under full fine-tuning, making it difficult for practitioners to trust third-party checkpoints. Existing defenses often rely on privileged assumptions (e.g., access to poisoned data or trigger/target knowledge), thereby limiting their applicability in realistic settings. We present DiSec (Disentanglement of potentially adversarial weights for Secure fine-tuning), a robust and label-efficient purification framework that uses only clean auxiliary text and does not rely on downstream supervision or attack signatures. DiSec elicits model-internal signals from this clean data to separate suspicious parameter components that are inconsistent with benign behavior, and then flags anomalous structures by jointly leveraging complementary spectral and generative views of outliers. Finally, DiSec performs a structure-preserving repair via layer-local prototype-based mean correction, yielding an idempotent update that depends only on non-adversarial statistics. Across diverse downstream classification tasks and PLM backdoor strategies, DiSec substantially suppresses attack success while preserving clean-task utility, offering a practical path to securing fully fine-tuned PLMs before deployment.
Citation
Das, S. (2026). Localizing and Repairing Backdoor-Sensitive Layers in Pre-trained Language Models for Secure Fine-Tuning. Graduate Theses and Dissertations Retrieved from https://scholarworks.uark.edu/etd/6343