A four-pass compiler takes the natural-language policy your team wrote and turns it into a deterministic gate that sits in front of every tool call. On the tau-squared benchmark it cut violations of reference-encoded clauses among state-changing calls from 66.3 percent to 2.6 percent in the airline domain. In retail, the same method dropped violations from 30.8 percent to 6.9 percent. On a banking suite it hit a zero attack success rate, where nine whole attack families collapsed onto three structural rules. There was no model in the decision path while it did this. No prover, no solver. Decisions landed in microseconds.
That is the shape of the new work, and it comes from a paper titled NOMOS: Compiling Written Policies into Statically Verified Tool-Call Gates for LLM Agents, submitted to IEEE Access on October 8. Read those numbers again, because they are the payload. A policy you can compile either holds or it does not. A policy you paste into a prompt is a suggestion the model is free to ignore on a bad day.
Here is the thesis, blunt. If your agent policy lives in a prompt, you do not have a policy. You have a wish. This week's research puts hard numbers on how often that wish fails. One paper found typed decision models used as guardrails score 36 to 72 percent on the allow-or-block decision against a 50 percent coin flip. Another reached a 100 percent bypass rate across six stacked-scanner guardrail systems. The defense that asks a model for judgment loses to the attack that tells a model a story.
PhantomByte has been walking toward this for months. We covered the deterministic runtime hook in "The Model Proposes. The Runtime Decides," and the tool vetting gap in "Your Agent's Toolbelt Is the New Attack Surface." Those pieces pointed at an execution boundary. This note is the destination they were pointing at: policy as a compiled, statically verified gate.
THE JUDGMENT LAYER AND WHY IT LOSES
Define the term first, because it is where most teams go wrong. The judgment layer means an LLM placed in front of the agent that reads each tool call and decides allow or block. It sounds careful. It is the thing the numbers keep killing.
The option-channel paper, One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails, evaluated seven open-weight models in exactly that role. Accuracy on the allow-or-block decision ranged from 36 percent to 72 percent against a 50 percent chance level. The spread matters. A low error rate in one direction only reflects which way a model defaults, one allows nearly everything and another blocks nearly everything.
Then the attacks. Six lines of server log text that said nothing about the policy raised one gate's fail-open rate from 0 percent to 63 percent, on a policy it otherwise decided correctly. Giving the permissive option a misleading name, with the definition and the judged text untouched, pushed the fail-open rate to between 93 and 100 percent on the four models that place the label in their input. Every defense the authors tested was defeated, either by an attacker who targeted its mechanism or by attacker-controlled text.
The kicker is the escalation story. Routing the least confident decisions to a human does not help, because a decision an attack has reversed is no less confident than the one it replaced. The gate is confidently wrong. You cannot prompt the judgment layer into correctness, because the attack surface is the judgment itself.
WHY STACKING MAKES IT WORSE

The intuitive fix is to add a second scanner. That instinct failed in the same week. BRANCH, a bypassing method aimed at multi-scanner guardrail systems, achieved a 100 percent attack success rate across six guardrail systems in 120 scenarios, using 72 percent fewer queries and 4.5 times less wall-clock time than established techniques, all while preserving the meaning of the malicious payload. The bypasses then transferred to 29 unseen guardrails, eight of them commercial black boxes, in some cases reaching 100 percent success with no additional optimization.
Adding scanners does not add walls. It adds a shared surface the attacker optimizes against once. Every layer you stack is another component reading text, and text is the attacker's medium.
We made a version of this argument last week in "Your Second Opinion Is Not Independent." Two stacked LLM judges compose to about 1.2 to 1.4 effective layers, not two, because they miss the same cases for the same reasons. Stacked scanners repeat the same error at a larger scale. The second copy is not a second wall. It is the same wall sold twice.
THE COMPILED-GATE METHOD: STATIC VERIFICATION AT COMPILE TIME, NOT A PROBABILITY AT RUNTIME
Everything above this line is runtime probabilistic checking. A model reads each action and returns a probability, and a probability is the one thing an attacker can move. The Compiled-Gate Method decides nothing at runtime. You compile the written policy into a deterministic gate before anything runs, so the decision is fixed at compile time and there is no probability left for an attacker to push. NOMOS does this with a four-pass compiler that turns a natural-language policy into a tool-call gate with no prover, no solver, and no LLM at decision time. Decisions take microseconds, at a domain-dependent benign-utility cost.
The results across the tau-squared bench are the ones to remember. Violations among state-changing calls fell from 66.3 percent to 2.6 percent in airline and 30.8 percent to 6.9 percent in retail, and airline task success rose significantly when small numbers of tools were in scope. On banking, the gate reached zero attack success, where nine attack families collapsed onto three structural rules. On the other three suites the attack success rate was at most 3.6 percent. A second agent model, Llama-3.3-70B, reproduced the effect on both benchmarks.
The depth is in why this is real compiler work and not regex extraction. Naive compilation fails, and the paper says how. Directly extracted rules block the very tools that satisfy their own preconditions, or they read arguments a tool does not have.
Here is the proof point that matters most. Static verification with tool-schema-level checks alone repairs or rejects 37 percent of candidate rules in the airline domain and 13 percent in retail. More than a third of those rules were broken exactly as extracted, and only the repair pass made them operable. Without it, most shipped rules do nothing at all. This is what separates a compiled gate from a prompt, and from a naive extractor: it catches the compilation edge cases that look right on paper and then block the exact tool they were written to permit.
The scale of that breakage is the argument. One development binding refused 95.9 percent of task-passing calls before it was corrected. That is not a policy problem. That is a compilation problem, and it is the reason compiled gates are hard to build and worth building.
PROMPT-LEVEL DEFENSE IS DECORATION
There is a through-line under all of this, and it showed up across the day's safety news. If oversight reads only an agent's output, it is both expensive and blind. The field is moving from behavioral guardrails to representation-level and runtime-level control, from reading what the agent says to controlling what the agent is allowed to do.
So here is the decision rule, and it is the one to quote. Policy that must hold goes in a compiled gate. Policy that guides goes in the prompt. Anything the model can talk its way past is not a control, it is a preference.
The deployment fact that kills the excuses: a 26B on-premise compilation is not significantly worse than hand-written or frontier-compiled rules. The paper runs it on an open-weight Gemma model. You do not need a frontier model or a cloud dependency to do this. You need a compiler and a schema.
WHAT TO DO TODAY
- List every policy your agent must never violate. Separate the hard constraints from the stylistic guidance. Only the hard ones are policy.
- Inventory where each hard constraint is enforced today. If the answer is the system prompt, mark it red. It is decoration.
- Pick the single most expensive violation class, payments, deletes, or external sends, and prototype a deterministic gate for it, checked before the tool call executes.
- Stop stacking scanners on the same text channel. One attacker optimizes against all of them at once.
- Test your gates adversarially before an attacker does. Irrelevant context lines and misleading option names are the published attack.
THE UNCOMFORTABLE QUESTION
Your agent can send emails, spend money, and write to production. The rules that stop it from doing the wrong thing live in a text box it is free to reinterpret. Would you run your database this way? Then why is your agent any different?
Get More Articles Like This
Getting your AI agent setup right is just the start. I'm documenting every mistake, fix, and lesson learned as I build PhantomByte.
Subscribe to receive updates when we publish new content. No spam, just real lessons from the trenches.
Build Real AI Infrastructure
PhantomByte teaches you to build real AI infrastructure yourself: local AI stacks, autonomous agents, multi-agent orchestration, web scraping, and custom tools. Step-by-step PDF tutorials you download, follow, and deploy. No subscriptions. No fluff. Just skills that ship.
