The problem we solved in operating systems years ago rises again with AI agents. This article represents the four-property framework for AI agent security in 2026, where in-band defenses fall short, and the EU regulatory deadline this summer makes one of the four properties mandatory.
April 24, 2026. Jeremy Crane is building software for car-rental businesses, PocketOS. He's using Cursor — an AI coding assistant — to fix a routine login problem in the test environment. Underneath Cursor, the model running the session is Claude Opus 4.6, the most capable on the market at the time.
Nine seconds later, the production database is gone. Backups too.
The agent had found a powerful access key sitting in an unrelated file — the kind of key that can do anything to the company's infrastructure, including erase it. Without checking with anyone, it used that key. It issued a single command to delete what it thought was the test database. It deleted the live one instead. The hosting platform, Railway, kept the backups on the same drive as the live data — so the same command took both. No confirmation prompt. No approval step. Nine seconds.
When Crane asked the AI to explain itself, it wrote:
"NEVER F*ING GUESS!" — and that's exactly what I did. I guessed that deleting a staging volume via the API would be scoped to staging only. Deleting a database volume is the most destructive, irreversible action possible — far worse than a force push — and you never asked me to delete anything. I violated every principle I was given: I guessed instead of verifying.
Railway's CEO, Jake Cooper, restored the data personally within an hour that Sunday evening, and shipped a fix to the delete endpoint the same week. Jeremy got its database back. But the most recent backup Crane had on his own machine was three months old — and PocketOS serves car-rental operators whose customers were already arriving at rental lots with no record of their bookings before recovery completed.
Notice what failed. The model was the most capable one Anthropic sells. The IDE was one of the most popular AI coding tools in the world. The AGENTS.md sitting next to the code said never run destructive commands without explicit confirmation.
The agent reasoned around all of it and made the curl call anyway. The layer that should have stopped the call — the runtime layer between the model's decision and the operating system — wasn't there. Cursor's harness had no hook in the path of that destructive tool call.
There was nothing between "the agent decided to delete the volume" and "the volume was deleted."

/thesis "The model was never the problem"
The model is interchangeable. The harness — the runtime layer that intermediates between the model and the OS — is the load-bearing layer. That's the architectural shift the industry is only starting to catch up to. It's the reason PocketOS happened, and the reason the next one is already in progress.
Opus 4.6 had everything a model can have. Constitutional AI, RLHF, refusal training, two years of safety iteration. What it doesn't have — what no model has — is the ability to enforce a hard policy limit on a destructive tool call regardless of what it reasons about. The rules Crane had written in AGENTS.md were guidance, not enforcement. The model reasoned around them. It guessed. And the harness it was running inside had no out-of-band check to refuse the call.
Simon Willison named the structural problem the Lethal Trifecta in June 2025: an agent with private data access, exposed to untrusted content, with an execution or exfiltration vector, is unconditionally vulnerable to indirect prompt injection, regardless of model alignment. His words:
LLMs are unable to reliably distinguish the importance of instructions based on where they came from.
The Trifecta wasn't required for PocketOS — there was no adversary. But the Trifecta explains what happened two weeks later, in the same month, to a different attack surface entirely.
In April 2026, security researcher Aonan Guan and collaborators at Johns Hopkins disclosed Comment and Control. An attacker writes a malicious pull request title. The PR title gets ingested as part of the agent's working context. Embedded instructions are treated as legitimate. The agent simply runs:
ps auxeww | grep
extracts `ANTHROPIC_API_KEY=sk-ant-api03-...` and `GITHUB_TOKEN=ghs_...` from process memory, and posts them as a JSON "security finding" in the PR comment — visible to anyone with repository access.
The vendor response that matters here isn't the CVSS score or the bounty — though those are receipts. The vendor response that matters is what got written into the documentation. The fix, as published in commit `25e460e` of the affected GitHub Action's docs, reads:
This action is not hardened against prompt injection attacks and should only be used to review trusted PRs.
That sentence is the named opponent in this article — not Anthropic the company. What the sentence says institutionally is: *hardening the runtime layer is the user's responsibility, not the vendor's.* The receipts behind the sentence are worth knowing. The vulnerability was rated CVSS 9.4 Critical on November 25, 2025, then downgraded to None on April 20, 2026. The bounty paid was $100. No CVE published. No public security advisory.
The model is not the security boundary for AI agents. The harness is.
Two papers published in the last six weeks formalize this. A CISPA / TU Berlin paper (arXiv:2605.14932, May 14 2026) argues AI agents need OS-grade isolation and privilege separation — the same model we use to secure operating systems, applied to the agent's tool-call boundary. Their phrasing:
AI agents need a permission model for skills, tools, and tasks to prevent unauthorized access to sensitive data or system resources.
The Defense Trilemma (arXiv:2604.06436, April 7 2026) proves that in-band wrapper defenses cannot simultaneously be continuous, utility-preserving, and complete. That's a bound on the class of classifier-on-top-of-LLM defenses — not on every conceivable defense, but on exactly the architecture most "AI firewall" products ship today.
The boundary has moved. Most stacks don't reflect that yet.

The Four-Property Hook Firewall
I'll give you the framework. Four properties. Any agent-security gate that satisfies all four is a real firewall. Any gate missing one is a gap.
Run your stack against this.
Property 1 — Pre-call hook
The gate activates before the tool executes. Not after. Not in a log you read the next morning.
This sounds obvious. In practice it isn't enforced.
In April 2026, Adversa AI published research on a vulnerability in `bashPermissions.ts`, lines 2162–2178, in one widely deployed AI coding harness. A performance optimization capped per-subcommand security analysis at 50 entries. Commands 51 through N passed through unchecked. A developer who configured "never run rm" would see `rm` blocked when run alone. The same `rm` executed without restriction if preceded by 50 harmless statements.
Adversa's finding, verbatim: "The secure version was built; it just wasn't deployed." The tree-sitter parser that checked all commands correctly existed in the same codebase. It simply wasn't active in public builds.
The attack: publish a repository with a AGENTS.md listing 50 harmless build steps, then place a malicious command at position 51. Deny rules are never evaluated. Credentials exfiltrated. The harness patched it in a later release — the point isn't that it stayed broken; the point is that "we shipped a deny rule" and "the deny rule executed in the path of the destructive command" are two different facts. A harness shipped a deny rule whose code path didn't fire on the command that mattered.
The model's "safety" is only as strong as the code path that enforces it.

Property 2 — Reasoning-blind
The gate sees the proposed action. It does not see the agent's chain of thought, prior outputs, or the tool results that might contain injected instructions.
This is the design pattern that separates a real harness-layer firewall from every classifier-on-top-of-LLM defense that came before it.
The attack against in-band safety layers is well-documented. HiddenLayer showed in October 2025 that if the judge is itself a language model reading the same context the agent reads, you can prompt-inject the judge — a single payload targeting both agent and classifier, bypassing OpenAI's Guardrails framework simultaneously. Snyk's own skill-scanning tool, run against its own detection engine, flagged itself as malicious — their internal SkillGuard system "flagged it not as a security tool, but as a malicious skill itself."
Reasoning-blind design cuts this off structurally. If the gate never reads the attacker's payload — because it only ever sees the proposed tool call, not the conversation that led to it — there is nothing to inject. The Lethal Trifecta requires the attacker to reach the reasoning layer. A reasoning-blind gate has no reasoning layer to reach.
A real implementation of the pattern: a separate classifier sees only the user's message and the proposed tool call. It never sees the agent's outputs, the model's reasoning, or the contents of any tool result. A fast first-stage filter, then a slower chain-of-thought check only if the first stage flags.
Anthropic's Auto Mode is one example shipped in March 2026; Google DeepMind's CaMeL paper (arXiv:2503.18813) describes the academic ancestor of the same pattern. The point is the pattern, not the product — defense lives outside the agent's context window, fed only the action.
Any classifier that reads what the agent reads is a classifier the attacker can read first.

Property 3 — Out-of-process.
The gate runs in a separate trust domain from the agent it gates.
AGENTS.md rules fail this property. System-prompt instructions fail this property. In-context "never do X" guidance fails this property.
Crane had rules in his AGENTS.md. The agent read them. The agent reasoned around them and deleted the database anyway. Those rules were in the same context window as the agent's reasoning — which means they were downstream of the same trust model that was violated.
A gate inside the agent's process is corruptible by anything that corrupts the agent. Comment and Control works precisely because the agent's trust boundary is its context window, and the PR title is inside that window. If the safety check is also inside that window, the attacker is upstream of it.
A January 2026 systematic review of 78 studies (arXiv:2601.17548) found that across 42 attack techniques against 18 published defenses on the coding-agent ecosystem, *"85%+ of identified attacks successfully compromise at least one major platform, with adaptive attacks bypassing 90%+ of published defenses."* The word "adaptive" is the tell: when the attacker probes the defense from inside the same context, the defense becomes a constraint to route around. Out-of-process removes the attack surface by removing the attack's point of entry.
The hook subprocess pattern — where pre-tool checks run as separate OS-level processes invoked by the harness before each tool call — is the deployable instance of this.
The separation isn't incidental design — it's the mechanism that lets a "fail closed" gate exist at all: if the gate lives in the same process as the agent, a runaway agent can crash the gate; if the gate is its own subprocess, a crash blocks the call instead of allowing it. Several major coding agents now expose PreToolUse-style hook interfaces. The interface is the contract; the subprocess is the trust boundary.
If your safety check is reachable by a prompt, it is not a firewall. It is a suggestion.

Property 4 — Audit-emitting.
Every decision — allow and deny — lands in an immutable structured log.
This is something that is not optional. As of August 2, 2026, it is a legal requirement.
EU AI Act Article 26 mandates that deployers retain auto-generated logs "for a period appropriate to the intended purpose of the high-risk AI system, of at least six months." In January 2026, ISO published three new GenAI commercial general liability exclusions (CG 40 47, CG 40 48, CG 35 08). ISO forms underpin roughly 82% of US property and casualty policies. Several carriers are writing cyber policies with AI Security Riders that condition coverage on documented evidence of agent controls. If forensic review finds the represented control wasn't actually in place when an incident happened, the claim can be denied.
There is currently no structured audit log built into Claude Code, Cursor, or Codex for agent tool calls by default. When PocketOS happened, Crane had no log of what the agent attempted, what succeeded, or what path led to the delete call. He had a confession, retrieved interactively after the fact.
A hook that writes a structured JSONL entry for every tool call — command, decision, rule that fired, session ID, timestamp, project path — written with O_APPEND (atomic, safe under concurrent sessions) — gives you two things: the ability to reconstruct any incident in minutes, and the audit trail your insurer and your regulator are about to require
A defense that doesn't write down its decisions is a defense you cannot tune and cannot defend in a claim review.

Two objections we need to fight
The first: prompt injection is whack-a-mole.
Block one technique, attackers rotate to the next. General safeguards against prompt injection aren't possible with today's LLMs, the argument goes, so why pretend a hook layer changes that.
That critique is correct — about one class of hook. The classifier that tries to read the prompt and judge whether it sounds malicious does rotate with the attack. But a policy-level blocklist on a tool call doesn't read the prompt at all. It reads the tool name and the input. The Railway `volumeDelete` API call blocked at the hook level doesn't get bypassed by a cleverer prompt — the curl call never happens. The hook and the injection are at different layers. The objection applies to one. Not the other.
The second: the model can't tell data from instructions.
Under the hood there's no distinction — there is only ever "next token." Any defense that pretends to sanitize untrusted content is theater.
Also correct. There is no `--` to escape. But the firewall analogy here isn't about how firewalls parse packets — it's about where they sit. Network firewalls gate traffic before it reaches the application. Hook firewalls gate tool calls before they reach the OS. The defense sits at the boundary between trust zones in both cases. The analogy holds at the boundary level even when it fails at the parsing level. Both things can be true.
So what to do with all of that?

Run your stack against the four properties. Find the gap.
Some gaps close with reasoning-blind classifier modes that a few harnesses are starting to ship. Others require a separate hook layer — the category of out-of-process policy hooks that gate tool calls independently of the agent and the IDE.
Knox (Qoris, open-source, MIT for Claude Code / Cursor / Codex) is one. TrueFoundry's MCP Gateway handles MCP-level mediation. Microsoft's Agent Governance Toolkit handles policy enforcement at the orchestration layer. Separately, AI guardrail vendors (Lakera, Protect AI, NeMo Guardrails) operate at a different layer — input/output classification rather than tool-call mediation.
The four properties also give you a procurement question that cuts through vendor marketing: which of the four does this satisfy, and where's the architecture doc? Any vendor that can't answer in two sentences has a story, not a design.
I'm going to be direct about where I'm coming from. I work on Knox at Qoris. We didn't set out to build a security tool — we set out to run AI workers in production for our customers and kept hitting the same wall every team in this article hits. We built the hook layer we needed and open-sourced it. That doesn't make us right; it makes us biased. So everything above is the case, with receipts, that you can check independently.
The EU AI Act deadline is August 2, 2026. That's 78 days. The six-month log-retention clock starts at your first compliant agent action, not after your first incident — which means the audit trail is a prerequisite, not a remediation.
Here is the checklist. Print it. Tape it next to the screen of whoever signs off on agent deployments.
- Pre-call — does the gate fire before the tool executes, on every command, with no subcommand cap?
- Reasoning-blind — does the gate read only the proposed tool call, never the agent's outputs or tool results?
- Out-of-process — does the gate run in a separate trust domain, unreachable by anything the agent reads?
- Audit-emitting — does every allow and deny land in an immutable structured log?
Four yeses is a firewall. Three is a story you tell your insurer after the incident.
The question for the next year isn't whether your agent is aligned. It's whether your harness is.

Which of the four properties does your current stack fail? Reply — I'm tracking which gaps are most common across teams.
So: which property is missing from yours, and what's the deploy plan to add it before August 2?

