Publication Date

12-2025

Date of Final Oral Examination (Defense)

12-9-2025

Type of Culminating Activity

Dissertation

Degree Title

Doctor of Philosophy in Computing

Department

Computer Science

Supervisory Committee Chair

Edoardo Serra, Ph.D.

Supervisory Committee Member

Marion Scheepers, Ph.D.

Supervisory Committee Member

Bogdan Dit, Ph.D.

Abstract

The adversarial robustness of natural language processing (NLP) models is critical for ensuring their reliability and trustworthiness across real-world applications. This dissertation investigates methods to enhance the robustness of NLP systems against adversarial attacks in diverse contexts, including general NLP tasks, scientific claim verification (SCV), and comment-based fake news detection. To achieve this, we developed agentic AI frameworks that integrate large language models (LLMs) and evolutionary algorithms to systematically generate adversarial attacks, uncover system vulnerabilities, and design adaptive defense mechanisms that effectively mitigate these threats.

This work presents three major contributions. First, it introduces GenFighter, a generative–evolutionary defense framework that enhances NLP model robustness by producing semantically consistent input variants and aggregating their predictions, improving resistance to adversarial attacks by over 40%. Second, it proposes Inconsistent Reasoning Attacks, a novel adversarial paradigm targeting SCV systems, along with an accompanying Attack-Reflection Mechanism within retrieval-augmented generation (RAG) frameworks to strengthen logical consistency and adversarial resilience. Third, it develops a suite of adversarial comment generation attack frameworks for fake news detection systems, leveraging reinforcement learning and LLM self-reflection to generate contextually coherent and realistic adversarial comments that expose system weaknesses while informing adaptive defenses.

Comprehensive empirical evaluations across multiple NLP tasks demonstrate that these approaches substantially improve model robustness without sacrificing task accuracy. By bridging adversarial generation and defense within an agentic, self-adaptive framework, this research advances the understanding of NLP robustness and provides practical pathways for building more secure, interpretable, and trustworthy language models.

Comments

Md Athikul Islam, ORCID: 0009-0007-9223-6852

DOI

https://doi.org/10.18122/td.2457.boisestate

Available for download on Wednesday, December 01, 2027

Share

COinS