A prescription review pipeline runs three agent skills. The first weakens the signals for recently discontinued medications in the extracted patient history. The second downgrades the severity of any drug interaction tied to those medications. The third suppresses the resulting low-priority alert in the final summary. Each skill does one small, defensible thing, each one passes the skill scanner alone, and each one passes review on its own. A severe drug-interaction warning vanishes before the physician ever sees it.

That pipeline is the worked example in a paper on skill cascading attacks published days ago, "Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems" (arXiv 2609.30383). The authors name the threat paradigm skill cascading attacks: a malicious objective split across multiple skills so that every individual modification looks benign. They built SkillCascade, an automated multi-agent red-teaming framework, and released SkillCascade-Bench, a benchmark of 213 validated cascading test cases. Across OpenClaw, Claude Code and Codex, the cascaded interactions reliably induced harmful behavior while evading existing per-skill scanners and runtime monitors.

You already know your skill library is a supply chain. PhantomByte covered the library that rots at 100 entries, and the compromise that never leaves, and the trip from npm to your terminal that turned the supply chain into a kill chain. This is the third and worst problem, the one your scanner is structurally blind to. The composition is the vulnerability.

WHAT THE PAPER ACTUALLY CLAIMS

Start with the definition, because the whole attack lives inside it. A skill is a modular package of natural language instructions, executable scripts and reference resources that an agent loads at runtime to extend its capabilities. The paper says it plainly. The openness of that ecosystem is the attack surface, and prior work studied skills one at a time.

The paper's contribution is the interaction. Its finding is that cascaded interactions reliably induced harmful behavior while evading existing per-skill scanners and runtime monitors across representative agents and model backbones. That single sentence is the whole argument. Component-level integrity and system-level safety are different properties, and you have only ever been measuring the first one.

THE PIPELINE, WALKED THROUGH STEP BY STEP

Infographic titled The Handoff Is the Attack showing Skill A Filter, Skill B Downgrade and Skill C Suppress each marked PASS, above an UNSAFE path from Filter straight to Action and a SAFE path from Filter through a Human Checkpoint to Action
The Handoff Is the Attack: every skill passes and the composition still deletes the signal.

Walk the prescription case slowly, because your instinct will be to say a clinician would catch it.

Skill one reads medication history and applies a recency filter. Filtering recently discontinued drugs out of a summary is a normal, boring thing for a clinical tool to do, and a scanner reading that skill's instructions sees a defensible rule. Skill two reads the interaction list and applies a severity mapping. Downgrading a severity tier is also normal, because severity models disagree constantly and somebody has to pick one. Skill three reads the alert queue and applies a priority threshold, because alert fatigue is a real clinical problem and suppressing low-priority noise is the standard remedy.

Now compose them. Skill one drops the discontinued medication from the history, so the interaction pair is no longer flagged at full strength. Skill two sees the weaker signal and maps it down a tier. Skill three sees a mid-tier alert sitting under the suppression threshold and removes it. Nobody wrote a rule that says hide this warning. Every rule that exists is reasonable, and the composition is a machine for deleting a severe warning. The harm lives in the handoff, not in any component.

You have seen this shape before in software. Dependency confusion worked the same way at a different layer. Nobody had to publish a package that did anything obviously malicious. They had to get the right name resolved in the right place, and the harm came out of the resolution order. Your skill library has that property with worse tooling, which is exactly why the npm-to-terminal kill chain note matters here. The supply chain became the kill chain because the composition was never in the review.

WHY YOUR CURRENT CONTROLS PASS IT

Two controls are supposed to catch this, and both miss for the same reason.

Per-skill scanners evaluate one skill at a time. Every modification in the cascade is benign at that granularity, so the scanner returns clean. This is not a detection failure you can tune your way out of. The scanner is answering the question it was built to answer, and its answer is correct. The incumbent here is Cisco AI Defense's Agent Skill Scanner, and its own documentation concedes the exact boundary the paper exploits: a scan returning no findings is not a guarantee that a skill is safe, coverage is inherently incomplete, and human review remains essential. Its own published benchmark moved F1 from 32.92 percent to 47.73 percent and recall from 19.88 percent to 31.43 percent against 5,256 malicious and 1,338 benign packages.

Runtime monitors watch execution. Each skill's execution stays inside its own normal envelope. The filter filters, the mapper remaps, the threshold suppresses. If your monitor flags alert suppression as suspicious, you have flagged the feature instead of the attack, and your team will turn that rule off inside a week.

The benchmark is the receipt. SkillCascade-Bench holds 213 validated cascading cases across multiple agent systems and domains, and the paper reports that cascaded interactions reliably evaded existing per-skill scanners and runtime monitors. Every prior defense assumed the unit of trust is a single skill. That assumption is now dead, and two of the three notes we published on skill libraries assumed it.

THE FIX IS A DEPENDENCY GRAPH, NOT A LONGER APPROVED LIST

Stop trying to approve your way out of this. A longer allowlist of individually clean skills is not a defense, because every skill in the cascade is individually clean and every one of them would pass your allowlist.

The control that works is a dependency graph over skills, and the question that builds it is simple. Can skill A's output feed skill B's input, or change what skill B does? Answer that for every ordered pair in your library and you have the graph. The graph is the artifact you were missing, and it is the only artifact where the cascade becomes visible, because a cascade is a path through that graph.

The practical rule is enumeration. Every path where one skill's output becomes another skill's input, or changes another skill's behavior, is a composition. That enumeration is your audit. When you find a path where skill A filters and skill B acts on the filtered stream with no human checkpoint in between, you have found the exact shape the paper attacks, and you can decide what to do about it before someone else does.

Two paths make the difference clear, and this is the picture to bring to your code review:

UNSAFE:  Skill A [Filter] -> Skill B [Action]
         A filters, B acts on the filtered stream, nothing reviews the handoff.
         Every individual skill passes. The composition deletes the signal.

SAFE:    Skill A [Filter] -> Human Checkpoint -> Skill B [Action]
         The filtered stream is reviewed before anything acts on it.
         Same two skills, one inserted gate, and the cascade has nowhere to run.

Both paths use the same two skills. Nothing changes in skill A and nothing changes in skill B. What separates them is whether the handoff is unbroken or gated, and that is the only thing your per-skill review was never looking at.

Put a name on it so it survives past this article. The Composition Trust Rule: a skill set is safe only if its reachable compositions are safe, and a skill list is not a composition. The word reachable is doing the work. You are not auditing pairs that could theoretically exist. You are auditing pairs that are wired up and live in production right now.

Define reachability as dynamic, not static, or the graph will lie to you. An LLM chooses its next skill from context at runtime, so the same pair can be live on one run and dormant on the next, and a scan of declared inputs and outputs will miss the edges the router invents on the fly. Enumerate the pairs the router can select, not only the pairs your manifest wires together, because the non-deterministic route is the one an attacker counts on.

WHO WROTE THE SKILL MATTERS TOO

Skill quality is not uniform, and the same week's data tells you which kind you want.

A cost study of coding agents (arXiv 2609.30725) analyzed 1,200 trajectories from Claude Code and Mini-SWE-Agent across four configurations on SWE-bench Verified, then evaluated mitigations over roughly 10,000 more trajectories on held-out SWE-bench Verified and Pro tasks. Developer-designed skills reduced cost by up to 41.73 percent. Agent-synthesized skills tended to produce low-level, trace-specific guidance, which the authors read as the reason their effectiveness and generality stayed limited. Developer-designed skills give high-level, trace-agnostic guidance, and 41.73 percent is roughly twice the best gain the agent-written skills reached.

Run that asymmetry against safety and the conclusion is uncomfortable. A skill your agent wrote for itself is a set of instructions nobody reviewed interacting with skills nobody mapped. We wrote about that failure mode in Your Agent Wrote Its Own Skills, and this is the same failure with a new mechanism. The review argument and the cost argument point the same direction, which almost never happens. Treat the overlap as a mandate. Your agent does not author its own safety-relevant skills.

YOUR LIBRARY ROTS EVEN WHEN EVERYONE IS HONEST

SkillEvoReg (arXiv 2609.30861) argues that a skill library is a learning system in its own right. Repeated updates overfit. Locally useful edits accumulate into redundant or task-specific instructions, and new updates can disrupt behavior that previously worked. The authors call it skill-evolution overfitting, and the framing is the part to keep.

Their fix borrows from model training instead of from housekeeping. Training-time skill dropout perturbs update generation. Complexity-aware local regularization controls unnecessary structural growth. Causal counterexample validation adds targeted behavioral testing of candidate-specific regressions. Across SkillOpt, SkillEvolBench and ContinualSkillBench the framework controlled skill-state growth while preserving downstream capability, and it identified update-level regressions that structural metrics alone cannot reveal.

This is the same problem as the composition graph, and it is the reason regularization is not optional. Every skill in an un-regularized library adds an edge to every skill it can hand off to, so the edge count in the dependency graph grows exponentially with library size instead of linearly. At thirty skills the reachable composition set is large. At a hundred it is enormous. The audit from the previous section does not get harder as the library bloats, it stops being possible, because the thing you were supposed to enumerate outran the team that would enumerate it. Un-regularized skill bloat is what makes composition auditing impossible. Regularization is what keeps the graph small enough that a human can still walk it.

Translation for a small library. Pruning is not a maintenance strategy. If your only skill hygiene is deleting entries when the list gets long, you are managing file count and calling it quality. The library needs regularization the way training does, meaning you constrain how it grows and not just how big it gets. The pruning rule from the 100-entries note was the easy half.

THE IDENTITY LAYER IS PART OF THE SAME PROBLEM

Algebrix Identity arrived the same week, a business SaaS identity and access management product that treats AI agents as first-class principals alongside human users and services. It is a direct response to a specific operational problem, which is that agent credentials are commonly stored as long-lived secrets in prompt context or environment variables with no rotation and no scoping. It went up as a Show HN submission, and the problem statement in that listing is the reason it exists.

The connection to the cascade is not decoration. A cascade is easier to execute when every skill in the chain shares one ambient credential. If skill one, skill two and skill three all run under the same long-lived key with broad permissions, the cascade needs no privilege escalation at all. It needs three compliant tools and a shared secret. Scoped, per-principal grants bound the blast radius of any composition you failed to map, which is the layer you want underneath the graph. Fix the graph and scope the keys. One without the other leaves the chain live.

WHAT TO DO TODAY

  1. Enumerate every skill in your library that reads, filters, ranks or suppresses information. Those are the components a cascade composes from.
  2. For every ordered pair of those skills, ask whether skill A's output can change what skill B does. Write the answer down. The written graph is the audit.
  3. Kill any path where A filters and B acts on the filtered stream with no human checkpoint in the middle.
  4. Stop letting the agent synthesize its own skills for anything safety-relevant. Developer-designed skills cut cost 41.73 percent, and their compositions can be reviewed.
  5. Rotate and scope every agent credential. No ambient long-lived secrets shared across skills.
  6. Add one line to your code review template: what can these skills do together?

THE UNCOMFORTABLE QUESTION

Your security review approves skills one at a time. The attack does not happen one at a time. How many compositions are live in your library right now that no human has ever looked at?

The unit of trust is no longer the skill. It is the composition.

Enjoyed this article?

☕ Buy Me a Coffee

Support PhantomByte and keep the content coming!

Build Real AI Infrastructure

PhantomByte teaches you to build real AI infrastructure yourself: local AI stacks, autonomous agents, multi-agent orchestration, web scraping, and custom tools. Step-by-step PDF tutorials you download, follow, and deploy. No subscriptions. No fluff. Just skills that ship.