Publication Date
12-2025
Date of Final Oral Examination (Defense)
12-9-2025
Type of Culminating Activity
Dissertation
Degree Title
Doctor of Philosophy in Computing
Department
Computer Science
Supervisory Committee Chair
Edoardo Serra, Ph.D.
Supervisory Committee Member
Marion Scheepers, Ph.D.
Supervisory Committee Member
Bogdan Dit, Ph.D.
Abstract
The adversarial robustness of natural language processing (NLP) models is critical for ensuring their reliability and trustworthiness across real-world applications. This dissertation investigates methods to enhance the robustness of NLP systems against adversarial attacks in diverse contexts, including general NLP tasks, scientific claim verification (SCV), and comment-based fake news detection. To achieve this, we developed agentic AI frameworks that integrate large language models (LLMs) and evolutionary algorithms to systematically generate adversarial attacks, uncover system vulnerabilities, and design adaptive defense mechanisms that effectively mitigate these threats.
This work presents three major contributions. First, it introduces GenFighter, a generative–evolutionary defense framework that enhances NLP model robustness by producing semantically consistent input variants and aggregating their predictions, improving resistance to adversarial attacks by over 40%. Second, it proposes Inconsistent Reasoning Attacks, a novel adversarial paradigm targeting SCV systems, along with an accompanying Attack-Reflection Mechanism within retrieval-augmented generation (RAG) frameworks to strengthen logical consistency and adversarial resilience. Third, it develops a suite of adversarial comment generation attack frameworks for fake news detection systems, leveraging reinforcement learning and LLM self-reflection to generate contextually coherent and realistic adversarial comments that expose system weaknesses while informing adaptive defenses.
Comprehensive empirical evaluations across multiple NLP tasks demonstrate that these approaches substantially improve model robustness without sacrificing task accuracy. By bridging adversarial generation and defense within an agentic, self-adaptive framework, this research advances the understanding of NLP robustness and provides practical pathways for building more secure, interpretable, and trustworthy language models.
DOI
https://doi.org/10.18122/td.2457.boisestate
Recommended Citation
Islam, Md Athikul, "Enhancing Adversarial Robustness in Natural Language Processing (NLP) Tasks Using Language Models" (2025). Boise State University Theses and Dissertations. 2457.
https://doi.org/10.18122/td.2457.boisestate
Comments
Md Athikul Islam, ORCID: 0009-0007-9223-6852