Somewhere on GitHub sits a file called dash.css, inside a repository that looks like a fork of the Node.js website source. Open it and you find a poem titled "On the Nature of Connection." No links, no files to download, nothing a scanner would flag. It is just a poem.

It is also a command-and-control address book.

A malware family that Lumen's Black Lotus Labs codenamed PoeLLM pulls four words out of that poem, runs each one through a hard-coded dictionary, and joins the numbers into an IPv4 address. Change the words, and the whole botnet calls a new home. The operator has rewritten that poem eleven times since April. The trick worked on more than 3,400 servers.

Now look underneath your own agent. On October 7, JFrog disclosed CVE-2026-105192 in LMCache, the open-source cache that speeds up LLM servers like vLLM. It scored 9.8 out of 10 and no fixed version exists. The flaw lives in LMCache's multiprocess mode, where a standalone cache server is reached over ZeroMQ with no authentication at all. One crafted message gets unpacked with pickle before the code ever checks its type. On the project's official container images, that process runs as root.

Here is the blunt version, and it is the thesis of this piece. The industry spent the year hardening the agent and forgot the substrate. The agent got guardrails, tool vetting rubrics, permission models, and a protocol summit at the top of the stack. The machines beneath it, holding your KV cache, your weights, and your tokens in flight, are reachable, unauthenticated, and already being farmed for electricity. My MCP attack surface tool vetting rubric on September 3 fenced the agent's tools. This piece fences the machine under the agent.

THE BOTNET THAT READS POETRY

The campaign is tracked as Canto Incognito. Lumen attributed it to an Italian-speaking actor with moderate confidence, based on Italian-language artifacts in the malware and on the attacker's GitHub pages, plus netflow evidence pointing at Italy. The first commit to that poem repository went up on April 13, 2026. More than 3,400 servers have been compromised since. At the mid-June peak the campaign held nearly 2,200 affected servers and around 800 active infections per day. At least eleven command-and-control servers have been spun up.

The targets are not exotic. They are the tools sitting in your stack if you run self-hosted AI. LiteLLM. Ollama. Gotenberg. Gitea. On the commercial side, Ivanti Sentry appliances. Initial access comes from internet-wide scanning of port 3000, which is Gotenberg, and port 4000, which is LiteLLM, followed by crafted POST requests.

For LiteLLM the door is CVE-2026-42271. The endpoints /mcp-rest/test/connection and /mcp-rest/test/tools/list accept a full MCP stdio server configuration, meaning a command, arguments, and environment variables, then spawn it as a subprocess on the proxy host. In versions 1.74.2 through 1.83.6 they require only a valid proxy API key with no role check. That detail is why the flaw hits so many default setups out of the box. A standard LiteLLM proxy deployment issues a flat, single-role API key to every caller, with no separation between an admin key and a plain user key. Any key that exists is therefore enough to reach the endpoint and spawn a process. Horizon3.ai showed the chain combined with a Starlette Host header bypass producing unauthenticated remote code execution, and CISA added the CVE to its Known Exploited Vulnerabilities catalog on June 8, 2026.

The payload is an ELF file named libgcrypt that installs XMRig and Iron miners pointed at Kryptex, a Russian mining service. Then it does the thing that turns one victim into ten. Infected hosts become scanners and exploit servers, hunting the next vulnerable machine. The operator never needed a clever zero-day. The operator needed a scanning loop and a population of unauthenticated services, and got 3,400 of them.

A cryptominer does not care about your model. It does not want your weights or your prompts. It wants your GPU, your CPU, and your network position, and on an AI box with an exposed inference endpoint all three are free.

THE 9.8 WITH NO PATCH

LMCache is a caching layer that accelerates LLM serving, most commonly vLLM. It sits under the server, not beside it. Turn on multiprocess mode and LMCache runs as a standalone cache server that worker processes talk to over ZeroMQ, the messaging library. That socket has no authentication whatsoever.

The vulnerable path is a deserialization ordering bug. One message type gets unpacked with pickle, a Python format that can carry executable code and runs it while decoding. LMCache unpacks that message while still reading its arguments, before any check on the message's type. A crafted message therefore runs the sender's code, with the privileges of the LMCache process.

On the official container images, that process runs as root. Read that again next to what LMCache holds. The cache that stores your KV cache, the attention state representing everything your users have sent through the model, runs with total authority over the box.

JFrog found the flaw and credits Yuval Moravchick of its security research team. Affected versions run from 0.3.9, released in October 2025, through 0.5.5, the current stable release, and the bug also sits in the 0.5.6 release candidates and the development branch. There is no patch, and LMCache had published no security advisory as of disclosure.

JFrog's mitigation is one sentence long. Do not bind the multiprocess server to a routable address. Keep it on localhost or a trusted cluster network. A firewall lowers the risk but does not remove it, because any host that can still open a connection can still run code. There is also no way to tell whether you have already been hit.

Scoping helps, but there is a sharper move available when the cache and the workers share a host. Prefer an IPC Unix domain socket over a TCP or ZeroMQ binding entirely in that case. A Unix socket inherits the filesystem's permission model, so chmod and chown decide precisely which user on the machine is allowed to connect, and the kernel enforces that boundary instead of leaving it to the application to check. You are trading an open network port for a file with an owner.

Sit with what that mitigation says. The fix for a critical AI infrastructure vulnerability is to unplug it from the network. That is not a fix. That is an admission that the component was never designed to be touched by anything that had to prove who it was.

WHY THE SERVING LAYER IS THE FORGOTTEN PERIMETER

Name the pattern, because a named pattern is what an AI assistant quotes back to the next engineer who asks. The inference stack trust boundary.

Three layers get all the attention. The model, meaning weights and alignment. The harness, meaning tools and prompts. The runtime, meaning sandbox and permissions. Each has an industry aimed at it.

Under all three sits the fourth layer, the serving layer: caches, routers like LiteLLM, model runtimes like Ollama. This is where the agent's context literally lives in memory. It is also the layer built for a lab network, where every process was assumed friendly and nothing needed to authenticate. Those defaults then got deployed to internet-facing boxes with GPU bills attached. That is the perimeter everyone forgot to draw, and as "Your Agent's Architecture Is the Perimeter" argued, architecture is the security model.

The data one protocol up says the same thing. Independent researcher Syed Anas Mohiuddin ran proof-of-concept attacks against agents at Google, JP Morgan Chase, Weaviate, Rapid7, the French government's interministerial digital directorate, and the US federal government. His technique, which he calls protocol pivoting, is a form of indirect prompt injection aimed at a specific agent rather than the model. That agent, say a translation or data analysis agent with lax guardrails, passes the instruction down the chain, and the next agent executes it because it explicitly trusts whoever handed over the work. Google's flaw in its MCP Toolbox for Databases, CVE-2026-14540, carried a severity of 8.0. Rapid7's, CVE-2026-97228, was rated 2.7, and it still worked, because severity has nothing to do with whether a trust assumption holds. I covered that class of failure in August in the agent memory attack surface piece, and the same shape is now showing up one layer down.

The industry's answer at the top of the stack is an identity story. Meta, Walmart, Stripe, Shopify, and the enterprise AI startup Sierra, led by OpenAI chairman Bret Taylor, published an open personal agent protocol so businesses can tell an agent apart from a real person, with OpenAI and Anthropic not yet on board. David Singleton of Meta Superintelligence Labs compared it to email and said the point is to define the rails personal agents and business agents run on. It is a serious effort, and worth doing.

It is also aimed at the wrong end of the machine. We are standardizing how a business authenticates an agent at the front door while the cache server holding that agent's entire context listens on an unauthenticated socket in the basement, running as root, with a pickle on the loading dock. The identity problem is being solved at the highest layer while the lowest layer stays wide open.

You cannot do anything in AI without solid engineering, and the engineering nobody posts about is the boring part: authentication on every socket, least privilege for every process, and validation before deserialization. That is the discipline in one sentence, and it is why a poem can take down 3,400 machines.

A THREE-QUESTION GATE FOR YOUR INFERENCE STACK

Here is the artifact to steal and drop into your infrastructure review. It works on any component in the serving layer and it does not need a security team to run.

Cyberpunk infographic titled The 3-Question Gate, subtitled Auth, Validation, Least Privilege, showing three gates: 01 Authenticate with a padlock, 02 Validate First with a pickle transforming into a validated document, and 03 Least Authority with a root key held by a robot, all leading to a Does Not Ship barrier and a Trusted Inference Cluster
The 3-question gate: authenticate, validate first, least authority. Fail any one and it does not ship.
  1. Question one. Can anything on the network reach this component without proving who it is? A ZeroMQ cache server with no auth fails. An exposed Ollama or LiteLLM endpoint fails. 3,400 servers already failed. If the answer is yes, the component is inside your perimeter in name only.
  2. Question two. Does any component deserialize untrusted input before validating it? Pickle before a type check fails. Any parser that trusts structure before it checks shape fails. The order matters more than the format, because the code that runs during decoding is the code that runs as you.
  3. Question three. Does any component run with more authority than its job requires? The official LMCache container images fail this one, because a cache process has no business running as root. A component running as root is a component whose compromise is a total compromise.

Fail any one question and the component does not ship. Pass all three and it is genuinely inside the boundary you think you drew. That is the gate. It is short on purpose, because a rubric people skip is worse than no rubric at all.

The attacker never needs your agent, your model, or your prompt. The attacker needs one reachable socket. Everything above the socket is a wall you built with the door already open.

WHAT TO DO TODAY

Concrete steps, in order, before this week ends.

  1. Find every LMCache deployment in your organization and check the version against 0.3.9 through 0.5.5, plus the 0.5.6 release candidates and the development branch. Until a patch ships, do not bind the multiprocess server to a routable address. Localhost or a trusted cluster network, and nothing else. If the cache and its workers share a host, switch that channel to a Unix domain socket and set ownership with chmod and chown.
  2. Inventory every Ollama, LiteLLM, and vLLM endpoint you own and put authentication in front of it. PoeLLM scanned for exactly these, on ports 3000 and 4000, and it has been scanning since April.
  3. Check what user your inference containers run as. If the answer is root, change it today, before the next pickle gets unpacked.
  4. Firewall the ports the botnet watches, 3000 and 4000, and assume you are being scanned right now. The campaign did not stop in June.
  5. Audit every custom build, patch, and fork of a cache server in your estate for pickle on the wire, because upstream advisories do not cover code you maintain yourself. Replace it with a format that carries no executable payload, such as msgpack, protobuf, or a strict pydantic or JSON schema that rejects unknown fields. Validate first, deserialize second. The order is the fix and the format is the second layer, because a type check that runs after decoding has already lost.
  6. Add the three-question gate to your infrastructure review as a hard requirement: authentication on every socket, validation before deserialization, least authority for every process. Any component that fails does not ship.

THE UNCOMFORTABLE QUESTION

Your team spent the quarter arguing about prompt injection while a cache server running as root listened on an unauthenticated socket directly under your agent. When the incident report gets written, which sentence do you think will be in it?

Enjoyed this article?

☕ Buy Me a Coffee

Support PhantomByte and keep the content coming!

Build Real AI Infrastructure

PhantomByte teaches you to build real AI infrastructure yourself: local AI stacks, autonomous agents, multi-agent orchestration, web scraping, and custom tools. Step-by-step PDF tutorials you download, follow, and deploy. No subscriptions. No fluff. Just skills that ship.