The threat model in plain English

Assume the agent, and everything it reads, is untrusted. Prompts, web pages, tool arguments and effect fields can all be attacker-influenced. Your policy files, providers and host are trusted.

Blocked

Reduced but not removed

Out of scope

A compromised host or process, malicious dependencies, OS-level escapes, and physical or functional-safety systems. See what it does not do.

The full table, with mitigations and residual risk for twenty threats, is on GitHub.

← All lessons