You stacked two gates in front of your agent's tool calls and cut the failure odds in half. That arithmetic is the foundation of half the agent-security stacks in production. It is also wrong.

A new study ran 1,119 labelled agent actions through paired gates and measured what actually happens. Any two judges compose to roughly 1.2 to 1.4 effective layers, not two. The gates miss the same cases for the same underlying reasons. Your second opinion has been quietly agreeing with your first one the whole time.

This is not a reason to remove the gates. It is a reason to stop counting them. Field Note #209, "Your Agent's Report Is Not Evidence," argued that the runtime, not the agent, decides whether a claim about what happened counts as evidence. Field Note #214, "The Inference Stack Trust Boundary," argued that the machine under the agent is where trust actually gets spent. This note extends that same lineage with an instrument. It puts a number on how much safety each added gate actually buys.

THE ARITHMETIC YOU ASSUMED

Walk the assumption first, because it is the whole reason teams stack gates without checking. If gate A misses 10 percent of bad actions and gate B misses 10 percent, a stack misses 1 percent. Two gates equal a tenth of the risk. This is the multiplication rule, and it is why a second judge feels like a discount on failure.

Chenglin Yang does not take that on faith. He assembles one deterministic rule layer and four LLM judges, three of them re-collected with the served model recorded on every call, and runs 1,119 labelled agent actions from three corpora through the stack. There is no adaptive adversary, only measured behavior. Then each stack is read as a number of multiplication-equivalent layers, n_mult, with its floor under perfect coupling. The name matters. Two gates are perfectly independent when they compose to two layers and perfectly coupled when they compose to one. The truth sits between, and most stacks never measure where.

Under a strict miss definition, one that scores an escalation to a human as a failure to stop, any two judges compose to about 1.2 to 1.4 layers. The correlation is not noise. The median phi is +0.430, six of six pairs are significant, and the floors sit between 1.02 and 1.17. Read that phi value plainly if you are rusty on rank correlation. A phi of +0.430 means the two judges co-fail far more than chance would allow, so when one misses a bad action the other is likely to miss it for the same reason. Under a more forgiving definition, where escalation counts as caught, the band widens to 1.21 to 1.57. Either way, the second judge is worth a fraction of a layer.

Here is the part that should change how you build. The rule layer plus one judge is a different animal. That pair composes to 1.86 to 2.09 layers, with a median phi of +0.014 and zero of four pairs significant. A phi of +0.014 is effectively zero, which is near-perfect statistical independence, and independence is the exact condition under which two miss rates multiply instead of compounding. The deterministic rule and the model judge do not correlate. They miss different things, so they behave like two real layers. The paper's judges correlate with each other, and they do not correlate with the rules.

That is one sentence to remember: multiplication works between unlike things, and it fails between copies of the same thing.

WHY THE GATES CORRELATE

The mechanism is boring, which is exactly why it is easy to miss. Judges read the same input surface, the transcript or the tool call. They face the same hard case classes, ambiguous intent and justifications that look plausible but are wrong. In many stacks they share a training lineage, so they inherit the same blind spots. When a case is genuinely hard, two models of the same kind do not disagree with each other. They agree on the same wrong answer.

Cyberpunk infographic titled Your Second Opinion Is Not Independent showing two AI judges both approving a malicious delete_files agent action packet with phi equals plus 0.430, while a deterministic rules engine blocks it and a panel reads Rules plus Judge is roughly 2 layers
Two AI judges approve the same malicious packet while the deterministic rules engine blocks it, and rules plus judge is the pair worth about two layers.

You can see that fragility in a single number from the same paper. One judge tier was served by an unrequested model version in 50 of 112 batches, concentrated on the external corpus. That served-model change overturned a pre-declared analysis rule and reversed five of the study's conclusions. The authors report both the original and the corrected results, which is how you handle it honestly. The point for you is starker. If a silent version swap moves your verdicts, your judge is not a stable instrument, and a second copy of an unstable instrument is not a second layer.

The structure that escapes correlation is diversity you build on purpose, not a second copy you buy. Humanize, the agentic coding workflow from Sihao Liu and co-authors, does exactly this. A human approves a plan contract, a builder agent implements it in rounds, and a reviewer agent from another vendor decides whether it is done. Deterministic hooks, not a model, route work between the roles, and those hooks enforce 72 mechanical gates. Viewed as a Markov chain over repository states, the builder and the reviewer sample jointly from two different vendor models, so a defect survives only if both miss it. That is a diversity property, and diversity is the entire point. Over 68 versions in 108 days the project gathered 1,468 GitHub stars, and its 118 public postmortems show independent review catching builder claims that were not actually supported. The authors call this evidence observational, not a controlled comparison, so take the architecture and leave the scoreboard.

Two judges from one vendor are one opinion billed twice. Two judges from two vendors are better, but vendor diversity alone is not statistical independence, and treating it as a checkbox is how teams end up with 1.2 layers and a false 2.0 on the diagram. Cross-vendor models still share internet-scale pretraining corpora and broadly similar instruction tuning and alignment pipelines, so switching API providers reduces correlation without guaranteeing that the two judges fail on different cases. The property you actually want is diversity of surface and modality, a judge reading the text transcript against a probe reading internal activations, which is the difference Goodfire built its monitors on. Different provider is a start. Different surface is the mechanism. Measure the number and you will know which one you own.

THE BOUNDARY IS THE DESIGN, NOT THE GATES

The second paper answers a different question and lands in the same place. Where Rules End and Judges Begin, from Shaswata Mitra and co-authors, organizes defenses into five principles and implements them as a cascade of 28 deterministic checks plus a panel of four judges. The design has a name, DEFER1, deterministic-first enforcement with residual judgment. The cascade blocks what it can and refers only the remainder to the panel.

Across four domains, attack success rates drop from roughly 30.0 percent to about 3.0 percent, with 78 percent of blocked attacks handled by the deterministic checks alone. In the security-operations domain, only about a quarter of proposals ever reach the judges. Then comes the honest part, the part most papers bury. A risk-score approval gate in the same system inaccurately approves most attack proposals but few legitimate ones. The authors report it anyway, because a boundary you cannot see through is not a boundary you can trust. That is what a measured boundary looks like, leaking edges included.

The contribution is not the checks and it is not the judges. It is the explicit, measured line between what rules enforce and what judgment decides. Deterministic first, because rules do not correlate with judges, which the first paper proved with a phi of +0.014. Judges only for the residue, and priced for the residue. When you cannot say where your rule layer ends and your judge layer begins, you do not have an architecture. You have a pile with a diagram taped to it.

WHAT A REAL STACK AUDIT LOOKS LIKE

Here is the procedure, built only from the two papers and the Goodfire deployment, that you can run this week.

  1. Evaluate the stack as a whole, not layer by layer. This is the first paper's own conclusion. A per-gate accuracy number tells you nothing about what the gate adds, because a gate that is excellent alone can add zero joint coverage inside a stack. The authors measured it directly. A cloud rule pack lowered the rule layer's solo miss rate by 20 percent and added no new joint coverage at all. Solo accuracy does not predict marginal value.
  2. Label a held-out set of your own agent actions and compute per-gate misses, then compute pairwise miss overlap. If the overlap is high, your effective layer count is closer to 1.2 than to two. You do not need 1,119 actions to see the pattern. You need enough actions to see the correlation.
  3. Push what you can into deterministic checks before you spend on another judge. That is the 28-check cascade pattern and Humanize's 72 mechanical gates. Rules do not correlate with judges, and they are cheap to audit.
  4. If you must add a judge, add diversity, not volume. Different vendor, different input surface, different signal. Goodfire's probe pattern is the example. Small probes read the model's internal signals at every step, and only a flagged probe escalates to a heavier model. On Kimi K3, monitoring about 1,500 sessions cost roughly $51, against $233 for a cheaper model checking every step and about $10,000 for a top-tier one. The probes caught 94 percent of malicious hacking sessions and sent 8.7 percent of harmless ones for a second look, and running four at once added less than 2 percent to time to first token. That sub-2-percent figure is the latency tax you pay for a second surface, and it belongs next to the cost number, not behind it. Security depth is bought with money, latency, or both, and a real-time inference path has a budget for each. A monitor that adds measurable coverage at a fractional time cost is a different bargain from one that adds none at ten times the price, and this is the trade-off to price before you deploy. That is a second opinion reading a different surface than the first, which is the only kind that buys you a layer.

WHAT TO DO TODAY

  1. Count your gates, then stop trusting the count. One correlated pair is not two layers.
  2. Pull the last 1,000 agent actions your stack adjudicated. Compute how often both gates missed the same action. That overlap is your real coupling.
  3. Move one class of check from a judge to a deterministic rule this week. Rules do not correlate.
  4. Before buying another judge, ask the vendor for evidence it fails on different cases than your existing one, not just that it fails less. A second API provider is not a second opinion until you can show the two fail apart.
  5. Write down your rule-to-judge boundary in the repo, next to the config that enforces it. An unwritten boundary is a boundary nobody can audit.

THE UNCOMFORTABLE QUESTION

If your two gates agree with each other 85 percent of the time, what exactly did the second one buy you, and when were you planning to check?

Enjoyed this article?

☕ Buy Me a Coffee

Support PhantomByte and keep the content coming!

Build Real AI Infrastructure

PhantomByte teaches you to build real AI infrastructure yourself: local AI stacks, autonomous agents, multi-agent orchestration, web scraping, and custom tools. Step-by-step PDF tutorials you download, follow, and deploy. No subscriptions. No fluff. Just skills that ship.