Offensive security · for the agentic age
The AI Security Playbook
How AI systems actually get attacked - and what actually stops it. Written for the people who have to test one, defend one, or sign one off.
44 chapters~91k words49 CVEs mapped one author, Iaroslav Mezin
Currently tracking: September 2026what changedthe incident board
What are you here to do?
Break it
Assess an AI system you are authorized to test, from recon to the report.
Build & defend it
Ship AI features and agents without shipping a breach, then catch what gets through.
Govern & advise
Answer the board, the auditor, and the regulator with something defensible.
New to AI security? Start here with the basics - no prior background assumed - or follow one system end to end.
The one idea behind all of it
A language model reads its instructions and your data in the same channel, with nothing that enforces the line between them. So nearly every vulnerability in this book is a trust-boundary error: text from an untrusted zone treated as a command in a trusted one.
The 90-second check follows from it. A system is exploitable for data theft when it has all three of these at once. Break any one leg and the path closes.
Private data
Can it reach data you would not publish?
Close it scope access on-behalf-of the user, just in time
Untrusted content
Does it read text an attacker can influence?
Close it quarantine or spotlight untrusted input
External comms
Can it send data out - mail, webhook, API?
Close it allowlist egress; gate irreversible actions
The one-page cheat sheet has the rest - the trust-boundary principle, the agentic stack, where the controls live.