One hundred and fifty production incident reports, gathered from open-source compound AI projects and anonymized enterprise deployments, were analyzed end to end. The failures sorted into 23 distinct modes. Almost none of them lived inside a model.

They lived at the boundaries. Between the retriever and the generator. Between the tool and the orchestrator. Between the agent and whatever it was wired into.

A compound AI system is multiple models, retrievers, and tools wired together, and the wire is where it breaks. Your eval suite scores each component in isolation. The taxonomy says that is exactly the wrong place to look. We already told you that verification is the new build system in "The Verification Bottleneck Is Your New Build System," and that an agent's own report is not evidence in "Your Agent's Report Is Not Evidence." This piece is what those two left implicit. This is the map of where the things you must verify physically happen.

The source is a paper by Rudrendu Kumar Paul and Sourav Nandy, accepted at the AIWILD Workshop at ICML 2026. It catalogs 23 failure modes across five categories, plus a resilience pattern catalog with effectiveness measured under controlled fault injection. Here is the framework it describes, the numbers it measured, and the audit you run on your own stack before the next incident review.

THE FIVE BOUNDARIES

Call this the Five Boundaries model. Every compound system has them, and every one of them is a place where two correct things can produce one wrong thing.

Infographic titled 5 Boundaries, 3 Failure Shapes, The Seams Are the System, showing a glowing AI core wired to five boundary modules labeled Retrieval, Generation, Tools, Orchestration, and Integration, with warning tags for low confidence retrieval, malformed structured data, race condition, external API changing shape, and corrupted results, above three panels labeled Cascade, Silent Degradation, and Coordination
Five boundaries and three failure shapes. The seams are the system.
  • Retrieval failures. The seam between the retriever and the generator. The shape it takes: the retriever returns nothing useful but returns it confidently, the generator receives a thin context, does not signal that the context is thin, and answers fluently anyway. Nothing errored, so nothing fired. Look at where a low-score retrieval result is allowed to proceed into generation without a hard gate.
  • Generation failures. The seam between the model and the format the rest of the system expects. The shape it takes: the model produces a valid-looking response that violates an implicit contract. A field that should be an enum arrives as free text. A missing value arrives as the string "null." The parser downstream accepts it because the shape is right. This is where schema drift and ungrounded claims live.
  • Tool failures. The seam between the tool and the orchestrator. The shape it takes: the tool returns a partial success, or a success-shaped error, and the orchestrator reads the shape instead of the content. The agent proceeds on a result that never completed. Look at your tool layer's error contract. If "it worked" and "it half worked" decode to the same object, you have a seam here.
  • Orchestration failures. The seam between steps in the plan. The shape it takes: step three depends on an assumption that step one silently invalidated. The orchestrator holds no model of its own preconditions, so it keeps executing a plan whose foundation is gone. This is also where retries, loops, and repeated actions chew through budget without converging. And once you run agents concurrently, this seam frays faster than anywhere else, because the failure is no longer about the plan at all. Two agents share a state store, Agent A reads a value that Agent B rewrote mid-turn, and nothing in the system serialized the two, so there was no locking mechanism to make the read and the write atomic. Neither agent errored. Each one acted correctly on a state that was true when it looked and false when it committed. Step invalidation is the sequential version of this problem, and state synchronization and race conditions are the parallel version, and the parallel version is the one that will not reproduce on your laptop.
  • Integration failures. The seam between your system and everything outside it. The shape it takes: an external API changes a field, a rate limit returns a status your client treats as transient, an auth token expires mid-run. Every component is behaving correctly against a world that moved. This is the category most teams never test, because it lives outside the repo.

That is five boundaries. One hour with your own architecture diagram will find at least three of them in your stack.

THE THREE FAILURE SHAPES

The paper's most citable finding is that boundary failures arrive in exactly three shapes. This is the part to memorize, because it explains why your current monitoring is structurally blind to them.

Cascading errors. One component's error becomes the next component's input. The first component logs a warning you already ignore. The second component produces garbage that still looks like a result. By the third, the origin is invisible. This is why root-causing an agent failure so often ends at "the data looked fine."

Silent quality degradation. The output stays plausible enough to evade standard monitoring while quality erodes. Your latency charts are green. Your error rate is zero. Your users are quietly getting worse answers. This is the failure your dashboards cannot see, and the paper measured a specific fix for it below.

Coordination failures. Every part behaves correctly and the collective result is wrong. This is the one people resist, because it violates the intuition that correct components compose into a correct system. They do not. Correctness is a per-component property. Wrongness can be a system property.

Here is why per-model evaluation is structurally blind to all three. A per-component eval tests a component against a frozen input distribution. The seam is precisely where that distribution is not frozen, because each component's input is produced live by the component before it. Every part passes its own eval while the seam fails. You are grading the actors individually and never watching the play.

THE GHOST IN THE LONG RUN

The sharpest instance of a seam failure is a seam in time.

A second paper, from XinPeng Shen and six co-authors, names a failure mode called GHOST: Governance Hazard from Overlooked Safety Constraints across Turns. Under entirely benign interaction conditions, a long-horizon agent executes an action that violates a safety constraint specified many turns earlier. The reported occurrence rate is 11.5 percent on GPT-5.5. That is not an edge case. That is a routine event, and the damage it describes is irreversible.

The theoretical result is the part that should change how you build. If the residual conditional violation hazard along each safe prefix is bounded below by a non-summable sequence, execution enters the hazard region almost surely. In plain terms: even when each individual turn looks safe, cumulative hazard still grows across turns. A safety check at session start does not bind turn forty.

The boundary here is between turn N and turn N plus one. Enforcement that is not re-evaluated at every action is enforcement that silently decays. The authors build a two-layer defense called STAR-Guard that couples restoration of historical semantic safety constraints with a deterministic pre-execution audit, and they report no GHOST events under their GPT-5.5 setup. The lesson is the two layers, not the name. Restore the constraint, then check it at the moment of action, every time.

SEEING THE SEAMS: TWO INSTRUMENTS

You cannot audit what you cannot see, and the honest reason most teams skip trajectory evaluation is cost. Two papers close that gap.

Trajectory evaluation you can actually afford. A third paper, from Linh-An Phan and five co-authors, presents LiteTrajEval. The observation underneath it is simple: not all raw tokens in a trace are equally useful for diagnosis, and modern agent traces are enormous. LiteTrajEval derives compact domain-specific rule profiles offline, preprocesses each trajectory online, marks heuristic failure signals, serializes the trace under a fixed global budget, and hands it to a single rubric-guided judge. Against AgentRx on public Magentic-One-style datasets, it improves failure-localization alignment with human annotations by roughly 20 to 35 percentage points, and up to 23 percentage points on tau-retail, while cutting cost by about 6 times and evaluation time by more than 8 times. It has already been deployed in an enterprise agentic platform, so this is not a lab result. The policy for you is blunt: an eval you can run every day beats a deep eval you run never.

Conditional failure knowledge. A fourth paper, from Changxiu Ji, Amy Lu, Qizheng Zhang, and Kunle Olukotun, presents Sentry, and it says something uncomfortable about how nearly every agent framework handles memory. Failure lessons are conditional knowledge. Kept permanently in the agent's context, they misfire when the relevant failure is absent, and removing them from an evolving playbook actually improves performance. Sentry instead runs alongside the agent, detects a failure, retrieves the matching lesson from an external store, verifies without access to task rewards whether the agent recovered, and stores a new lesson only if it did. The full playbook never enters the context. Across multiple agentic benchmarks, Sentry beats the strongest runtime-intervention baseline on every benchmark by 37 percent on average, and beats the strongest context-evolution baseline by 39 percent on the two benchmarks where both were evaluated. Their controlled experiments also show that exposing the full playbook to the agent lowers performance even when the relevant lessons stay available on demand. That last sentence is a direct rebuke to the "stuff everything into the system prompt" pattern, and it is measured, not asserted.

The distinction worth holding onto is the one between working memory and retrieved operational lessons, because conflating them is what produces the overload in the first place. Working memory is the context window, the scratch space an agent needs to reason about the task in front of it, and it is finite and should be kept small. Retrieved operational lessons are not context at all. They are execution guardrails, indexed by the failure they repair and pulled in only when that failure actually occurs. Stuffing the second category into the first does not make an agent wiser, it makes it more distracted, and a context padded with lessons that do not apply to the current task is a direct hallucination vector, because the model will reach for a nearby lesson and apply it to a situation it was never written for. Sentry's design treats memory as out-of-band infrastructure, conditionally exposed at the moment of need, which keeps the context window clean and turns hard-won failure knowledge into something closer to a guardrail than a background reading assignment.

THE SEAM AUDIT

Five boundaries, three shapes, one hour of mapping. Run this on your own stack this week.

First, list every boundary in your system. Component to component, tool to orchestrator, orchestrator to plan step, turn to turn. Write them as edges in a graph, not as a list of services. The edges are the system. The nodes are the parts you have been testing. Here is the shape of it, and the arrows are the blast radius:

[Retriever] --Seam 1--> [Generator] --Seam 2--> [Orchestrator] --Seam 3--> [Tool]

Second, for each boundary ask one question. If this seam corrupted its output right now, which monitor would fire? If the answer is none, you have found a silent-degradation seam, and it is invisible by construction. This question is worth more than any dashboard you will build this quarter.

Third, re-evaluate safety constraints at action time, not at session start. The GHOST result is unambiguous. A constraint checked once is a constraint that decays.

Fourth, move failure lessons out of the context window and into a store retrieved on failure. Conditional exposure, never permanent loading.

Fifth, budget trajectory evaluation as a fixed line item, not as an emergency purchase after an incident. Offline rule profiles plus a fixed serialization budget are what make it affordable enough to become routine.

WHAT TO DO TODAY

  • Pull your architecture diagram and mark every arrow. Each arrow is a seam.
  • For each arrow, name the owner. An unowned boundary is an undetected boundary.
  • Add a hard gate where a low-confidence retrieval result can reach generation, and where a malformed tool result can reach the orchestrator. Fail closed at the seam, not inside the component.
  • Write down your tool layer's error contract, then check whether partial success is distinguishable from success. If it is not, that is a one-line fix with a very large blast radius.
  • Instrument the turn-to-turn constraint re-check and log every re-evaluation, because a constraint you cannot prove fired is a constraint you cannot prove exists.
  • Put a number on your recovery path. The paper found that systems implementing three or more resilience patterns from its catalog cut mean-time-to-recovery by 71 percent compared with unstructured monitoring baselines. Its controlled fault injection measured the individual patterns too: circuit breakers reduce cascade propagation by 89 percent, output quality gates catch 73 percent of silent degradation before user impact, and component isolation reduces blast radius by 64 percent. Those are the numbers to put in front of whoever controls the reliability budget, because they are the difference between a reliability argument and a reliability opinion.

THE UNCOMFORTABLE QUESTION

Your last outage review named a component. The taxonomy says the component was innocent and the wire was guilty. How many of your boundaries have no owner?

Enjoyed this article?

☕ Buy Me a Coffee

Support PhantomByte and keep the content coming!

Build Real AI Infrastructure

PhantomByte teaches you to build real AI infrastructure yourself: local AI stacks, autonomous agents, multi-agent orchestration, web scraping, and custom tools. Step-by-step PDF tutorials you download, follow, and deploy. No subscriptions. No fluff. Just skills that ship.