Ask a model a question about a document that says nothing about it. It answers anyway. Fluently, confidently, with a paragraph cited that does not contain the answer. You have probably blamed the prompt, the model, or your retrieval pipeline. The defect is older than all three of them.

Softmax attention weights sum to one. That single property of the equation means every attention head must distribute exactly 100 percent of its attention somewhere, even when nothing in the context deserves any of it. There is no zero setting. The head cannot say nothing here is relevant, so it does the only thing the math allows. It averages whatever it has and hands the average back as though it meant something.

A paper posted to arXiv in September (2609.22005, Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention, by Richard Zhe Wang) names that gap, names a second one sitting next to it, and then shows that a small addition to the value pathway supplies both. It also shows something more useful for anyone running models in production: which of the two benefits you actually collect depends on how large your model is.

PhantomByte field note #198 covered sparse attention dropping the block that held your answer. This is the layer underneath it. Even when no block deserves your answer, attention hands you one regardless. That is not a tuning problem you can out-engineer at the prompt layer. It is a decision the architecture made before training started.

WHY SOFTMAX CANNOT ABSTAIN

An attention head is one of the parallel read heads inside a transformer layer. Each head scores how much every token in the context should matter to the token being computed, then mixes the context according to those scores.

Here is the mechanism, and it is short. Softmax forces the scores into a probability distribution. They must sum to one. Every head therefore spends its full budget of attention on every token it processes, forever. There is no way to spend zero. When a query arrives and every key in the context is irrelevant, the distribution does not collapse to nothing. It spreads thinly across tokens that have nothing to do with the question, and the head emits the resulting blend as its output.

That is where confident nonsense is manufactured. Not in the weights file, and not in the tokenizer. In the equation. The model is not failing to recognize that the context is empty. The model is structurally forbidden from representing that fact as an output.

This is the first missing primitive, and the paper gives it a name: abstention, the ability of a head to produce no output when nothing is relevant. Softmax attention has no abstention primitive. It has to answer.

WHY IT CANNOT FILTER EITHER

The second gap is a mirror of the first, and it lives on the receiving side. A head's output is a weighted average of value vectors, which means it passes interference through as faithfully as it passes signal. Garbage that arrives in the value pathway gets read, mixed, and forwarded with the same care as the token you actually wanted.

There is no internal organ that refuses to read what is poisoned. Define this one the way the paper does: noise filtering, the ability to suppress what a head reads instead of what a head says.

You have watched this happen. It is the reason retrieval-augmented systems degrade so sharply when the retrieved chunk is wrong, and it is the reason a single bad document in a long context can drag an otherwise correct answer sideways. Teams blame the retriever, rerank harder, and add guardrails to the prompt. The model still has no mechanism for rejecting the chunk once it is in the window. In my read of the results, this is the single most underrated failure mode in modern inference stacks, and it is architectural rather than procedural.

WHY THIS IS LETHAL IN AGENT LOOPS

Retrieval is the easy case. The hard case, and the one most of your production traffic actually runs through, is an agent deciding what to do next.

In an agent loop the context is not a document. It is a set of tool schemas, the execution history of everything already tried, the current state of the scratchpad, and the task itself. The model scores all of it and emits an action. Now apply the primitive gap to that decision. When no available tool fits, and when the correct move is to do nothing at all, the attention heads are still distributing a full budget of weight across the tool definitions sitting in the context. There is no setting that says none of these.

That is forced tool calling, and it is not a hallucination in the usual sense. It is the same forced average from the earlier sections, aimed at your tool registry. The model does not select a tool because a tool is right. It selects because the mechanism must produce a selection, and the only candidates in the window are tools. Hand an agent four tools and a request that needs none of them, and watch which one it grabs.

Scratchpad drift has the same root. Once a low-relevance action lands in the history, the filtering gap forwards it, and the next step reads its own bad move back as established context. Two missing primitives, one failure loop, and the loop grows more confident with every turn.

If you ship autonomous systems, this is the sentence to carry into the design review: agent tool-routing models are structurally forced to emit a selection even when no selection is correct.

THE VALUE GATE: ONE FIX, TWO PAYCHECKS

A value gate is a learned scalar applied to the value pathway. In plain terms, it lets a head scale down or zero out what it writes into the residual stream, instead of always passing an average through to the next layer.

An infographic titled Two Missing Primitives: Abstention plus Noise Filtering, showing irrelevant documents averaging through softmax attention into a Value Gate labeled Abstention and Filtering, with relevant signal mixed with noise in the value pathway output.
The two missing primitives funnel through a single value gate.

Wang reports that this one mechanism supplies both missing primitives in part, and that finding unifies a set of earlier results that had been credited to different causes. The interesting part is the ablation across matched models from 10M to 350M parameters. At 10M, the gain from gating is almost entirely abstention. By 350M, filtering contributes as much as abstention does. The two benefits are largely additive, with only a small overlap between them.

Read that as a decision rule, because it is one. Small models gain mostly the ability to shut up. Large models gain both capabilities, roughly half and half. The paper calls this out directly: what a study observes about gating depends on the scale it ran at. A result measured at 10M is an abstention result, and it will not transfer cleanly to a 70B deployment.

I am naming this rule the Abstention-to-Filtering Shift, so you can carry one sentence into a design review: as models grow, the payoff from gating the value pathway migrates from abstention toward noise filtering, and by roughly the 350M mark the two pay equally.

One architectural detail worth keeping, and it is the detail that explains a behavior you have seen a thousand times. A gate controlled only by each value leaves the attention sink in place. An attention sink is the habit of dumping leftover attention weight onto the first token of the sequence, usually the begin-of-sequence token. The sink exists because of the primitive gap itself. Heads that must spend their entire budget need somewhere harmless to put the weight they have no use for, so they park it on a token that carries almost no content. The sink is not a feature anyone designed. It is a workaround for not being able to spend zero.

Value gating and screening remove the need for the workaround. If a head can decline to write, or a threshold can reject an irrelevant key outright, there is no leftover budget to park and no null token required. A query-controlled gate removes the sink entirely, while a gate determined by each value alone leaves it standing. That is the cleanest test of which fix you are actually holding.

SCREENING: RELEVANCE ON A BOUNDED, ABSOLUTE SCALE

The gating work fixes the output side. A second paper attacks the score itself, and I think it is the more fundamental of the two.

Screening Is Enough (arXiv 2604.01178, by Ken M. Nakanishi) starts with a definition of absolute relevance. Relevance is absolute when its values sit on a fixed bounded scale, when they depend on neither the competing keys nor the sequence length, when they need no sequence-length-dependent calibration, and when all of them are allowed to be exactly zero.

Standard attention fails every clause of that definition. Its scores are relative, they shift with whatever else is in the window, and their scale changes as the context grows. Screening replaces that with an explicit threshold that turns bounded query-key similarities into relevance values. The consequences are exact rejection, empty selection, and direct inspection of relevance on one common scale.

Empty selection is the whole ballgame. It is the abstention primitive, installed at the relevance score rather than at the output, which means the head never has to average garbage in the first place because the garbage never enters the average.

The controlled comparison is what sells it. Twelve attention mechanisms, one matched Transformer backbone. Only screening held both low long-context perplexity and retrieval that kept working beyond the training context, and it did so with no inference-time scaling tricks. Multiscreen, the architecture built on top of it, runs parallel gated screening tiles and keeps those long-context gains while posting better parameter efficiency, stronger general zero-shot performance, lower training cost at larger scales, and lower model-side time to first token than Transformer baselines. Nakanishi also ships a normalization design that keeps Multiscreen training stable at a learning rate of 1, and an adapted version of it that stabilizes a standard Transformer at the same rate.

A model you can train at a learning rate of 1 without divergence is not a small result. That is the kind of detail that shows up in someone else's training loop six months later.

TENSOR-LEVEL ABSTENTION IS NOT AN SFT REFUSAL

The obvious objection is that you can train the behavior. Fine-tune on I do not know and no tool needed, add a refusal class, and the model says the right words. Plenty of teams stop there and call the gap closed.

Look at what that patch actually is. A refusal string is generated late, after the attention layers have already done their work. The attention operation produced its forced average, that average entered the residual stream, and the MLP layers then did their best to decode something sensible from a representation that was already contaminated. You taught the model to apologize for garbage it had no mechanism to refuse.

That distinction is the whole argument. Refusal is a learned output pattern sitting on top of a broken input. Abstention is a structural property of the attention operation itself, and it keeps the garbage out of the residual stream before the next layer ever reads it. A value gate or a screening threshold changes what the next layer reads. An SFT refusal changes only what the final layer writes.

Label this as my read, because it is one: the refusal string is a receipt, not a fix. It documents that the model noticed the problem after the fact. It does nothing to stop the problem, which is why the same failure returns the moment the context shifts in a way your fine-tuning set never covered.

WHAT THIS MEANS FOR YOUR STACK TODAY

Deployability first, honestly. These are research results. The gate ablations run on matched models between 10M and 350M parameters, with pretrained checks up to 20B. Screening is evaluated on a matched backbone, not on a frontier production model. Nobody is shipping this into your inference server tomorrow.

What you can act on right now is not the code. It is the framing, and the framing changes several things you are currently doing wrong.

Prompting cannot fix an architectural refusal to abstain. When you write only answer if the context is relevant into a system prompt, you are asking a mechanism that mathematically must produce output to behave as though it could produce none. Sometimes the model complies at the surface. The compliance is a learned pattern, not a capability, and it breaks under pressure. In my opinion the instruction is worse than useless, because it convinces the team the failure has been handled.

Test the irrelevant case directly. When you evaluate long-context vendors or candidate models, construct a context that contains no answer at all and score the refusal rate. Most teams only ever test retrieval with the answer sitting somewhere in the window. That test cannot detect a missing abstention primitive, because the failure only appears when there is nothing to find.

Check the scale before you believe an attention paper. A gating result at 10M may be an abstention result rather than a filtering result, and it will not transfer to a 70B deployment. This is now the first question I would ask about any attention ablation that crosses my desk.

WHAT TO DO TODAY

  • Add an answer is not in the context test set to your eval suite. Score refusals as a first-class metric, not just accuracy on answerable questions.
  • Stop filing confident answers to irrelevant contexts as a prompting bug. File it as architecture, because that is what it is.
  • Test your agent's tool routing with impossible requests. Hand it a task no available tool can serve and score whether it invents one anyway.
  • If you build or fine-tune small models, gating buys you abstention. If you run large models, expect filtering to matter just as much.
  • Read arXiv 2609.22005 for the gate ablations and arXiv 2604.01178 for the 12-mechanism screening comparison before your next long-context model selection.
  • When a vendor claims long-context reliability, ask how their attention handles a zero-relevance query. Watch how quickly the room goes quiet.

THE UNCOMFORTABLE QUESTION

Your model has never once been able to say nothing here. Not in any query you have ever run, on any context you have ever pasted in. So ask the question that follows from it: of the answers you shipped this month, how many were fluent averages of garbage the architecture was structurally forbidden to reject?

You cannot prompt your way out of a missing primitive. You can only build it in.

Enjoyed this article?

☕ Buy Me a Coffee

Support PhantomByte and keep the content coming!

Build Real AI Infrastructure

PhantomByte teaches you to build real AI infrastructure yourself: local AI stacks, autonomous agents, multi-agent orchestration, web scraping, and custom tools. Step-by-step PDF tutorials you download, follow, and deploy. No subscriptions. No fluff. Just skills that ship.