The threat model in plain English
Assume the agent, and everything it reads, is untrusted. Prompts, web pages, tool arguments and effect fields can all be attacker-influenced. Your policy files, providers and host are trusted.
Blocked
- Self-authorization: the agent granting itself authority or approval.
- Policy bypass via prose: "ignore the rules" in any text field changes nothing.
- Provider failure: errors, timeouts and malformed responses all mean
FAIL_CLOSED. - Stale authority: expiry required, freshness checked, authority re-checked just before the action.
- Effect swapping: changing the action after approval changes the digest, so the approval no longer applies.
- Replay: an effect or approval can be used once, including under concurrent races.
- Config tampering by malformed input: invalid policy never loads and never falls back to allow.
- Missing audit trail: no receipt, no effect.
Reduced but not removed
- Unwrapped paths: if code calls the raw function, nothing governs it. Use
senscheck auditand thegovernance-reviewskill, and un-export raw functions. - Receipt deletion: unsigned receipts in a place the agent can write can be removed.
- Multi-process replay: single-use tracking lives in one process's memory.
- A compromised human channel or host.
Out of scope
A compromised host or process, malicious dependencies, OS-level escapes, and physical or functional-safety systems. See what it does not do.
The full table, with mitigations and residual risk for twenty threats, is on GitHub.