Page
1
(This page has no text content)
Page
2
(This page has no text content)
Page
3
Adversarial Machine Learning
Page
4
(This page has no text content)
Page
5
Adversarial Machine Learning Mechanisms, Vulnerabilities, and Strategies for Trustworthy AI Jason Edwards San Antonio, TX, USA
Page
6
This edition first published 2026 © 2026 John Wiley & Sons Ltd All rights reserved, including rights for text and data mining and training of artificial intelligence technologies or similar technologies. No part of this publication may be reproduced, stored in a retrieval system, or transmitted, in any form or by any means, electronic, mechanical, photocopying, recording or otherwise, except as permitted by law. Advice on how to obtain permission to reuse material from this title is available at http://www.wiley.com/go/ permissions. The right of Jason Edwards to be identified as the author of this work has been asserted in accordance with law. Registered Offices John Wiley & Sons, Inc., 111 River Street, Hoboken, NJ 07030, USA John Wiley & Sons Ltd, New Era House, 8 Oldlands Way, Bognor Regis, West Sussex, PO22 9NQ, UK For details of our global editorial offices, customer services, and more information about Wiley products visit us at www.wiley.com. The manufacturer’s authorized representative according to the EU General Product Safety Regulation is Wiley-VCH GmbH, Boschstr. 12, 69469 Weinheim, Germany, e-mail: Product_Safety@wiley.com. Wiley also publishes its books in a variety of electronic formats and by print-on-demand. Some content that appears in standard print versions of this book may not be available in other formats. Trademarks: Wiley and the Wiley logo are trademarks or registered trademarks of John Wiley & Sons, Inc. and/or its affiliates in the United States and other countries and may not be used without written permission. All other trademarks are the property of their respective owners. John Wiley & Sons, Inc. is not associated with any product or vendor mentioned in this book. Limit of Liability/Disclaimer of Warranty While the publisher and the authors have used their best efforts in preparing this work, including a review of the content of the work, neither the publisher nor the authors make any representations or warranties with respect to the accuracy or completeness of the contents of this work and specifically disclaim all warranties, including without limitation any implied warranties of merchantability or fitness for a particular purpose. No warranty may be created or extended by sales representatives, written sales materials or promotional statements for this work. The fact that an organization, website, or product is referred to in this work as a citation and/or potential source of further information does not mean that the publisher and authors endorse the information or services the organization, website, or product may provide or recommendations it may make. This work is sold with the understanding that the publisher is not engaged in rendering professional services. The advice and strategies contained herein may not be suitable for your situation. You should consult with a specialist where appropriate. Further, readers should be aware that websites listed in this work may have changed or disappeared between when this work was written and when it is read. Neither the publisher nor authors shall be liable for any loss of profit or any other commercial damages, including but not limited to special, incidental, consequential, or other damages. Library of Congress Cataloging-in-Publication Data Applied for Hardback ISBN: 9781394402038 ePDF ISBN: 9781394402052 ePUB ISBN: 9781394402045 Cover Design: Wiley Cover Image: © sankai/Getty Images Set in 9.5/12.5pt STIXTwoText by Straive, Chennai, India
Page
7
To my beloved family. To my mom and dad, who are no longer with us but whose love and guidance continue to inspire me every day. To my wife, Selda, and my children, Michelle, Chris, Ceylin, and Mayra, for your endless love, patience, and encouragement. To my close family members: Derek, Meltem, Nilos, and Ken. And to my sisters, Robin, Kelly, and Lynn. This book is dedicated to you, with all my love and gratitude.
Page
8
(This page has no text content)
Page
9
vii Contents Preface xi Acknowledgments xiii From the Author xv Introduction xvii About the Companion Website xxi 1 The Age of Intelligent Threats 1 The Rise of AI as a Security Target 1 Fragility in Intelligent Systems 3 Categories of AI: Predictive, Generative, and Agentic 5 Milestones in Adversarial Vulnerability 8 Intelligence as an Attack Multiplier 10 Why This Book and Who It’s For 12 Recommendations 14 Conclusion 16 Key Concepts 16 2 Anatomy of AI Systems and Their Attack Surfaces 21 The Architecture of Predictive, Generative, and Agentic AI 21 The AI Development Lifecycle: From Data to Deployment 24 Classical Machine Learning vs. Modern AI Pipelines 26 Identifying Entry Points: Training, Inference, and Supply Chain 28 Security Debt in the Model Development Lifecycle 31 Recommendations 33 Conclusion 35 Key Concepts 35 3 The Adversary’s Playbook 39 Threat Actors: Profiles, Motivations, and Objectives 39 White-Box Attack Techniques and Methodologies 41 Black-Box Attack Techniques and Methodologies 44 Gray-Box Attack Techniques and Methodologies 47 Operationalizing AI Attacks: Tactical Methodologies and Execution 49 Advanced Multi-Stage and Coordinated AI Attacks 52 Recommendations 54 Conclusion 55 Key Concepts 56
Page
10
viii Contents 4 Evasion Attacks—Tricking AI Models at Inference 61 Core Principles and Mechanisms of Evasion Attacks 61 Gradient-Based Evasion Techniques 64 Linguistic and Textual Evasion Methods 67 Image- and Vision-Based Evasion Techniques 69 Evasion Attacks on Time-Series and Sequential Models 72 Recommendations 74 Conclusion 76 Key Concepts 76 5 Poisoning Attacks—Compromising AI Systems During Training 81 Fundamentals and Mechanisms of Training-Time Poisoning 81 Label Manipulation and Clean-Label Poisoning Techniques 84 Backdoor and Trojan Insertion in Training Data 86 Poisoning Attacks on Federated and Distributed Learning Systems 89 Poisoning Attacks Against Reinforcement Learning (RL) Systems 91 Poisoning Attacks on Transfer Learning and Fine-Tuning Processes 94 Recommendations 96 Conclusion 98 Key Concepts 98 6 Privacy Attacks—Extracting Secrets from AI Models 103 Core Mechanisms and Objectives of AI Privacy Attacks 103 Membership Inference Techniques 106 Model Inversion Attacks and Data Reconstruction 109 Attribute and Property Inference Attacks 111 Model Extraction and Functionality Reconstruction 114 Exploiting Privacy Leakage Through Prompting Generative AI 117 Recommendations 119 Conclusion 120 Key Concepts 121 7 Backdoor and Trojan Attacks—Embedding Hidden Behaviors in AI Models 125 Fundamental Concepts of AI Backdoors and Trojans 125 Backdoor Trigger Design and Optimization 128 Data Poisoning Methods for Backdoor Embedding 130 Trojan Attacks in Transfer and Fine-Tuning Scenarios 132 Embedding Backdoors in Federated and Decentralized Training 135 Advanced Trigger Embedding in Generative and Agentic AI Models 137 Recommendations 140 Conclusion 141 Key Concepts 142 8 The Generative AI Attack Surface 147 Architectural Foundations of Large Language Models 147 How Generative Architectures Expand Attack Opportunities 150 Exploiting Fine-Tuning as an Adversarial Vector 152 Prompt Engineering as an Adversarial Exploitation Pathway 155
Page
11
Contents ix Technical Risks in Retrieval-Augmented Generation Systems 157 Leveraging Model Internals for Generative AI Exploitation 160 Recommendations 163 Conclusion 164 Key Concepts 165 9 Prompt Injection and Jailbreak Techniques 169 Technical Foundations of Prompt Injection Attacks 169 Direct Prompt Injection Methods and Input Crafting 173 Indirect Prompt Injection via External or Retrieved Content 175 Jailbreak Techniques and Semantic Boundary Exploitation 177 Token-Level and Embedding Space Manipulations 180 Contextual and Conversational Injection Strategies 182 Recommendations 185 Conclusion 186 Key Concepts 187 10 Data Leakage and Model Hallucination 191 Technical Mechanisms of Data Leakage in Generative Models 191 Membership and Attribute Inference via Generative Outputs 195 Model Inversion and Training Data Reconstruction 197 Hallucination Exploitation in Generative Outputs 199 Prompt-Based Extraction of Memorized Data 202 Exploiting Multi-Modal and Cross-Modal Leakage in Generative Models 204 Recommendations 207 Conclusion 208 Key Concepts 209 11 Adversarial Fine-Tuning and Model Reprogramming 213 Technical Foundations of Adversarial Fine-Tuning 213 Semantic Perturbation Methods for Adversarial Fine-Tuning 216 Embedding Covert Behaviors via Adversarial Prompt Conditioning 219 Advanced Trojan Embedding via Fine-Tuning Gradients 221 Cross-Model and Transferable Adversarial Fine-Tuning Attacks 223 Model Reprogramming via Adversarial Fine-Tuning Techniques 226 Recommendations 228 Conclusion 229 Key Concepts 230 12 Agentic AI and Autonomous Threat Loops 235 Technical Foundations of Agentic AI Systems 235 Technical Manipulation of Autonomous Decision Loops 238 Exploitation of Agentic Memory and Context Management 241 Agentic Tool Integration and External API Exploitation 244 Technical Embedding of Autonomous Chain Injection 246 Exploitation of Environmental Interactions and Stateful Vulnerabilities 248 Recommendations 251 Conclusion 252 Key Concepts 253
Page
12
x Contents 13 Securing the AI Supply Chain 257 Technical Mechanisms of Supply Chain Poisoning in AI Models 257 Artifact and Model Checkpoint Contamination Techniques 260 Technical Exploitation of Third-Party AI Libraries and Frameworks 263 Dataset Provenance and Annotation Manipulation Techniques 265 Technical Exploitation of Hosted and Cloud-based Model Infrastructure 268 Artifact Repositories and Model Zoo Contamination Methods 270 Recommendations 272 Conclusion 273 Key Concepts 274 14 Evaluating AI Robustness and Response Strategies 277 Technical Foundations of AI Robustness Evaluation 277 Metrics for Evaluating AI Security and Robustness 279 Robust Optimization Methods and Adversarial Training 282 Certified Robustness and Formal Verification Techniques 285 Technical Benchmarking Tools and Evaluation Frameworks 287 Technical Analysis of Robustness Across Model Architectures and Modalities 289 Recommendations 292 Conclusion 293 Key Concepts 294 15 Building Trustworthy AI by Design 299 Technical Foundations of Security-by-Design in AI Systems 299 Robust Embedding and Representation Learning Methods 302 Technical Approaches to Adversarially Robust Architectures 304 Technical Integration of Formal Verification in Model Design 306 Technical Frameworks for Runtime Anomaly Detection and Filtering 308 Technical Embedding of Model Interpretability and Transparency 310 Recommendations 313 Conclusion 315 Key Concepts 315 16 Looking Ahead—Security in the Era of Intelligent Agents 319 Technical Foundations of Future Agentic AI Systems 319 Emerging Technical Attack Vectors in Agentic Systems 322 Technical Exploitation of Multi-Modal and Cross-Domain Agentic Capabilities 325 Future Technical Capabilities in Automated Adversarial Generation 327 Technical Mechanisms for Evaluating Advanced Agentic Robustness 330 Technical Embedding of Ethical Constraints and Safety Mechanisms 332 Recommendations 335 Conclusion 337 Key Concepts 337 Glossary 341 Index 367
Page
13
xi Preface Artificial intelligence has moved rapidly from research projects to systems that make decisions in healthcare, finance, defense, and daily life. With this growth comes a sobering reality: intelli- gent systems are vulnerable. They can be manipulated, deceived, or subverted in ways that tradi- tional security practices were never designed to address. That reality is what inspired me to write this book. For more than two decades, I have worked in cybersecurity, and in recent years, I have focused much of my effort on education—both in the classroom at several universities and through BareMetalCyber.com, where I develop resources for learners and professionals alike. Across all of these settings, I have seen a growing demand for practical guidance on how to secure AI systems, not just how to build or apply them. Students, engineers, analysts, and executives all ask the same core questions: How do these attacks work? What risks do they pose? And what can we do to defend against them? This book is written to answer those questions. It is intended for a wide range of readers— engineers and scientists designing AI models, security professionals tasked with defending them, students preparing to enter the field, and decision-makers who must evaluate risk and policy. Rather than simply cataloging threats, the chapters are designed to build a framework for under- standing how adversaries operate and how defenders must respond. Throughout, I have included real-world examples, case studies, and exercises to make the material as relevant and actionable as possible. The work I do with learners, both through Bare Metal Cyber and in academic programs, has shown me how vital it is to bridge theory with practice. That is the spirit in which this book is offered. My hope is that it not only equips readers with tools and knowledge but also encourages a culture of responsibility, resilience, and trust in the systems we are building for the future. Dr. Jason EdwardsOctober 30th, 2025 New Braunfels, TX, USA
Page
14
(This page has no text content)
Page
15
xiii Acknowledgments I am deeply grateful to all those who have contributed to the creation of this book. First and foremost, I would like to express my heartfelt appreciation to my family for their unwavering support and understanding throughout the writing process. To my wife, Selda, and my children, Michelle, Chris, Ceylin, and Mayra, your patience and encouragement have been my anchor. I am also thankful for the love and support from my extended family: Derek, Meltem, Nilos, and Ken, and my sisters Robin, Kelly, and Lynn. I am indebted to the organizations and the many fellow educators who have been pivotal in my professional development and the success of this book: Hallmark University, Moravian University, IronCircle, Cybrary, and the more than 100K LinkedIn subscribers who follow me. To the millions of readers and listeners who follow my cybersecurity content each year on BareMetalCyber.com, your trust and engagement are the fuel that keeps this work moving forward. Thank you for making this journey possible. I also encourage everyone reading this to support a cause close to my heart—the Beagle Freedom Project, which fights for the rights and rescue of animals used in laboratory testing. You can learn more about their work and my children’s cybersecurity book series, starring Darwin the Cyber Beagle, at CyberBeagle.kids. To find more of my books, check out CyberAuthor.Me. To join a community of Cybersecurity Professionals visit BareMetalCyber.com.
Page
16
(This page has no text content)
Page
17
xv From the Author I wrote this book because the moment demands it. Artificial intelligence is no longer an experimental frontier—it is embedded in healthcare, finance, transportation, government, and national defense. Yet, as machine learning systems take on increasingly consequential roles, many of them remain profoundly vulnerable to adversarial manipulation. These risks are not theoretical. They are happening now. Attackers are already deceiving, poisoning, extracting, and reprogramming AI systems in the wild. Unfortunately, most organizations—and even many developers—are unprepared. My goal with this book is to fill that gap. Defending AI in the Age of Intelligent Threats: Adversarial Machine Learning is not just a catalog of attack types or a manual for specialists. It is a call to awareness and a structured roadmap for defense. I wrote it to help engineers, researchers, security professionals, and decision-makers understand not only how AI systems fail under attack, but also why those failures occur and what can be done to mitigate them. This book presents a unified framework for considering adversarial machine learning across both predictive and generative systems, with a focus on attacker goals, capabilities, and lifecycle considerations. I also explore areas that are often under-addressed, such as prompt injection, backdoors in foundation models, jailbreaks, and supply chain risks in open-weight and hosted models. What sets this book apart is that it spans the spectrum—from evasions at inference time to train- ing data poisoning to insider threats in model fine-tuning pipelines. It speaks to developers and to architects, regulators, and risk managers who are trying to build trustworthy systems in an environ- ment that is rapidly shifting. I’ve done my best to make the material accessible without sacrificing rigor and to equip readers with a vocabulary, a set of mental models, and a shared language that mirrors emerging standards from NIST and others. We are entering an era in which AI will be a permanent part of critical infrastructure—and that means it will be a permanent attack surface. This book is my contribution to the growing effort to secure that future. I hope it serves as both a warning and a toolset. If it helps even a small group of practitioners design more robust systems, deploy more defensible models, or challenge the false sense of security that can follow a successful demo, then it has done its job. Dr. Jason EdwardsOctober 2025 New Braunfels, TX, USA
Page
18
(This page has no text content)
Page
19
xvii Introduction Artificial intelligence is no longer an experimental curiosity. It powers financial transactions, diagnoses medical conditions, navigates vehicles, and produces content that shapes public opinion. With each passing year, AI becomes more tightly interwoven with the systems we depend on for stability and trust. This integration makes the security of AI not a theoretical concern but a matter of practical necessity. When an AI system fails under attack, the consequences ripple outward—affecting not just organizations but entire societies. The unique fragility of AI systems stems from their very strength: the ability to generalize from data. Unlike traditional software, which follows explicit human-coded rules, AI learns patterns that may be imperceptible to human observers. Those patterns can be manipulated by adversaries in ways that bypass human intuition. A stop sign altered with a few pixels can fool an autonomous vehicle. A prompt injected with subtle instructions can bypass filters in a generative model. These failures are not random glitches; they are pathways for deliberate exploitation. The challenge is magnified by accessibility. Attack tools once restricted to academic labs are now widely available. Open-source models, pre-trained weights, and online tutorials put adversarial capabilities into the hands of anyone with modest resources. As adoption accelerates, defenders must recognize that the threat landscape is democratized. It is not only nation-states or organized crime groups that pose risks; individual attackers can and will probe AI systems at scale. Why AI Security Matters Now Recent years have demonstrated that adversarial attacks are not limited to research papers. Manip- ulated transaction data has bypassed fraud detection systems. Face recognition systems have been tricked with adversarial patches printed on glasses or clothing. Large language models have been coaxed into generating sensitive or harmful outputs through carefully engineered prompts. Each case underscores a hard truth: intelligent systems can be steered toward errors in ways that humans cannot reliably anticipate. These risks are amplified by the supply chain that underpins modern AI. Organizations frequently rely on pre-trained models, third-party datasets, and open-source frameworks. Each dependency widens the potential attack surface. A poisoned dataset can spread through countless downstream applications. A compromised model checkpoint uploaded to a public repository can silently embed vulnerabilities in systems around the world. Unlike traditional software flaws, these risks are baked into the learning process itself, making them harder to detect and remediate after deployment.
Page
20
xviii Introduction At the same time, competitive pressure drives organizations to prioritize speed and innovation over robustness. Models are deployed quickly, often with minimal adversarial testing. Validation tends to emphasize average-case accuracy while overlooking edge cases where attacks thrive. This creates a false sense of security: a system that performs well under normal conditions may collapse under adversarial stress. The gap between adoption and defense is widening, and closing it has become urgent. Scope of This Book This book is designed to cover the adversarial landscape across the full life cycle of AI systems. It begins with the foundations, explaining the categories of AI—predictive, generative, and agentic—and why each introduces distinct vulnerabilities. Predictive systems, used in classifica- tion and forecasting, face evasion, poisoning, and inference attacks. Generative systems, including large language models, face prompt injection, jailbreaks, hallucinations, and leakage risks. Agentic systems, capable of autonomy and tool use, present systemic risks that extend into the physical world. The middle sections of the book dig into the attacker’s playbook. They describe how adversaries think, what resources they need, and how they operationalize attacks across white-box, black-box, and gray-box conditions. Readers will explore attack types in detail—evasion at inference, poison- ing at training, backdoor embedding, privacy extraction, and the unique exploits of generative and agentic AI. Each attack is paired with real-world examples, illustrating not just the theory but the consequences of adversarial exploitation. The final sections address defense and governance. They survey technical approaches such as adversarial training, certified defenses, and runtime anomaly detection, as well as systemic approaches that integrate security into the development lifecycle. Beyond technology, they also consider governance frameworks, regulatory requirements, and the organizational culture shifts needed to prioritize AI security. The goal is to equip readers with not just a catalog of tools but a structured framework for evaluating risk and building resilience. Who This Book Is For This work is written for a broad but connected audience. Engineers and data scientists will find detailed technical breakdowns of attacks and defenses. Cybersecurity professionals will see how adversarial AI maps onto established risk models and where it requires new approaches. Leaders and policymakers will gain a vocabulary for discussing AI risk in terms that align with governance and compliance frameworks. In each case, the goal is to provide both accessibility and depth, meet- ing readers where they are while pushing them to expand their perspective. Students will find the text valuable as a bridge between academic research and practical appli- cation. Exercises, discussion questions, and key takeaways are included to reinforce learning and stimulate critical thinking. Instructors can adapt these features for coursework, while independent learners can use them for structured self-study. Whether used in the classroom, professional devel- opment, or individual exploration, the book is designed to be flexible and supportive of multiple learning paths. For executives and decision-makers, the content highlights strategic implications as well as tech- nical ones. Understanding adversarial AI is not only about preventing system failures; it is about