Skip to content

Offensive security · for the agentic age

The AI Security Playbook

How AI systems actually get attacked - and what actually stops it. Written for the people who have to test one, defend one, or sign one off.

44 chapters~91k words49 CVEs mapped one author, Iaroslav Mezin

Currently tracking: September 2026what changedthe incident board

A language model reads its instructions and your data in the same channel, with nothing that enforces the line between them. So nearly every vulnerability in this book is a trust-boundary error: text from an untrusted zone treated as a command in a trusted one.

The 90-second check follows from it. A system is exploitable for data theft when it has all three of these at once. Break any one leg and the path closes.

1

Private data

Can it reach data you would not publish?

Close it scope access on-behalf-of the user, just in time

2

Untrusted content

Does it read text an attacker can influence?

Close it quarantine or spotlight untrusted input

3

External comms

Can it send data out - mail, webhook, API?

Close it allowlist egress; gate irreversible actions

The one-page cheat sheet has the rest - the trust-boundary principle, the agentic stack, where the controls live.