ChatGPT-6 Astra just cleared the orc starting zone in World of Warcraft in 40 minutes with zero deaths, and it never saw a single frame. It navigated by parsing the game's raw server network traffic and by reading quest data out of the server's own SQL files. No screenshots. No UI grounding. No vision tokens. The developer behind the run, the maker of the open-source agent-wow client, put it plainly: the model was more than capable of working at the protocol layer.

Look closely at how that run was built, because the architecture is the story. The model did the planning, and deterministic code did the execution. C++ pathfinding and direct SQL reads handled everything with one correct answer, and the model was left free to decide what to do next. Nobody asked the language model to trace a route pixel by pixel or to generate a coordinate per token. It set the goal, a compiled helper computed the path, and the model read the result back as data. That is the orchestration versus execution split, and it is the reason the run finished instead of flailing.

Meanwhile your production agents are burning roughly five times the tokens of a human doing the same job, and the ratio is trending toward ten times. More than 85 percent of those agent tokens are cached prompt re-reads, according to the a16z chart built on OpenRouter data. The agent rereads what it already saw. Two numbers, one lesson. The most expensive thing you can give an agent is eyes. The cheapest is a socket.

If you have read my earlier notes on the KV cache, you know where those tokens go. "Your Agent's KV Cache Dies at Every Turn Boundary" and "Your Agent's KV Cache Is Eating Your Inference Budget. Here Is the Fix." covered the mechanics of the bill. This note is about how to stop generating the tokens in the first place, by choosing the control surface before you choose the model.

THE THREE CONTROL SURFACES

Every agent drives software through exactly one of three surfaces, and the choice is an engineering decision you make, not a property of the model.

Infographic titled The Three Control Surfaces comparing pixels (screenshots to clicks), the accessibility tree (structured UI to actions), and the raw protocol (packets to data to control), with a banner reading Choose the Control Surface Before the Model
The three control surfaces: pixels, the accessibility tree, and the raw protocol.
  1. Pixels are screenshots in and clicks out. This is the default because it works anywhere a human works, and it is the most expensive option available, because every frame is a token farm and every click is a guess at coordinates.
  2. The accessibility tree is the structured description of an interface that assistive technologies already read. It is cheap and stable and survives cosmetic redesigns, but on most platforms it is read-only, and it is blind to anything the UI does not expose.
  3. The raw protocol is REST endpoints, wire packets, and database reads. This is the surface that does not care what the screen looks like, because there is no screen in the loop.

There is also a hybrid worth naming on its own: the browser's own protocol layer. Chrome DevTools Protocol and the DOM accessibility tree let a web agent read a page and act on it without a single vision token. When no public REST API exists, CDP and DOM-tree driving are the protocol stand-in, and they are the actionable bridge for web agent developers who cannot get a backend socket this quarter. Call it protocol-adjacent and treat it as a first-class option, not a workaround.

The Astra run is a proof of surface three. The setup used agent-wow, an open-source AzerothCore client described by its repo as a WoW client designed for autonomous AI agent players, running against a private World of Warcraft 3.3.5a server. One prompt in Codex. The agent built a module that captured 28 types of server messages and held them in memory, and a Python script polled those messages to assemble its picture of the world and to send actions back into it. For quests it data-mined the server's SQL files for quest givers, turn-ins, and spawn points, which is roughly what a human does in hours on a fan wiki, except the server's own files are the data the server actually runs on. For navigation it built a C++ helper that plotted routes across AzerothCore's own navigation mesh files using the Detour pathfinding library, returning waypoints as coordinates or an error when no complete path existed. The model never computed a waypoint in its own head; it handed that job to a library built for it and consumed the answer as structured output. It planned in sequence, working prerequisite quest chains first, selling junk, equipping upgrades, training abilities before the final cave, and picking up both cave quests at once so it could finish them together. The developer noted that pathfinding is normally where heuristic bots fall apart, and called this agent's pathfinding optimal.

The failure mode of pixels is not cost alone. A pixel-reading agent cannot verify what it did. It infers state from an image and then infers success from another image, and nothing in that loop can contradict the inference. That is the same structural problem the empty commitment paper names, a promise of an action after the current turn that nothing in the agent's tools or runtime can carry out, where the emptiness follows from the configuration alone and no later trajectory is needed to detect it.

WHY PACKETS BEAT PIXELS

Start with the cost math. Agents consume about five times the tokens of human users and are heading for ten times as retrieval and tool loops expand, and the cause is mechanical. They reread. A screenshot multiplies that behavior, because the same interface gets re-encoded into a fresh image at every step, and the model pays to look at pixels it has already seen in a slightly different arrangement. A packet is smaller than the frame it renders. A query response is smaller than the table it draws.

Reliability runs the same direction. Pixels force the agent to infer state from an image, so every action depends on a perception step that can fail quietly, and a misread button is indistinguishable from a working one until something downstream breaks. The protocol hands the agent state as data, with field names and values and types, so the agent reads instead of squinting.

Auditability is where the gap becomes structural. A packet-level action stream is already structured and already loggable. When your agent acts on the protocol, its history is an audit trail it produced as a byproduct of doing the work. When it acts on pixels, your audit trail is a folder of PNGs and a hope that the pixels captured the relevant state.

Then add verification, which is the part most teams skip. The Rules to Tools paper shipped prepared executable checks of public scientific requirements to matched repair groups that shared written checks, starting programs, model, and budgets, and the group that also received a callable implementation of those checks went from 26 of 30 complete repairs to 29 of 30. On five development-exposed tasks with alternate starting programs, the same split ran 3 of 10 with text and 7 of 10 with tools. The finding is narrow and it is useful: converting a written rule into something a program can execute changes outcomes more than adding more explanation does. A protocol gives your agent something to check against. A screenshot gives it something to squint at, and squinting is not verification.

WHY TEAMS STILL CHOOSE PIXELS, AND WHAT IT COSTS THEM

An honest version of this argument has to answer the obvious objection. Plenty of teams automate through the interface not because they are naive, but because the interface is the only thing that holds still. Internal APIs are frequently undocumented, un-versioned, and prone to silent breaking changes. The UI, by contrast, is comparatively stable, and it changes on purpose, because a human is looking at it and someone in the business will notice the day it moves.

So the trade-off is real and it deserves a name. Schema drift is the quiet failure of protocol control, and it is a maintenance line item you have to budget for the same way you budget for the vision tokens. A socket you have to babysit is still usually cheaper than a screenshot you have to trust, but pretending the babysitting is free is how teams get burned and retreat to pixels permanently.

The answer is not to argue drift away. It is to instrument for it. Put contract tests on every endpoint your agent depends on. Validate the response schema at runtime and fail loudly when a field disappears instead of letting the agent improvise around a null. Version the interface, even informally, and keep a pinned copy of the last known-good contract. Then add one monitor whose whole job is to tell you the endpoint changed before your agent finds out the hard way. Do that, and drift becomes a known cost you control rather than an excuse for vision.

WHAT THIS LOOKS LIKE IN ENTERPRISE SOFTWARE

Most business software already has a socket. A REST API, an SQL interface, an MCP tool, an A2A endpoint. The typical enterprise agent is clicking through a graphical interface while an API that performs the same action sits unused two layers down.

Take invoice processing.

  • A pixel agent screenshots each queue screen, reads the invoice number out of an image, and clicks into a detail view, and it pays vision-token prices for that on every invoice.
  • A tree agent reads the queue's labels and gets the same fields more cheaply, until it hits a column the UI never exposed.
  • A protocol agent queries the queue's API and receives every field as data, including the ones the UI truncates.

Same job. Three cost curves. Three failure profiles. One rule resolves most of it: for every screen your agent looks at, ask what the screen is rendering, because the screen is always rendering data that came from somewhere you can query. If you have already settled which stack your data lives in, as I argued in "Three AI Stacks. Your Data Already Picked One.", then the control surface is the next decision you make and the cheapest one to revisit.

THE SCREEN-OR-SOCKET DECISION RUBRIC

Give the agent a screen only when one of two conditions holds.

  • First, no protocol exists and building one costs more than the vision tokens will. Legacy thick clients, Citrix farms, that one Java app from 2009 that nobody wants to touch.
  • Second, the task's ground truth is visual by definition. Whether a rendered layout looks broken has no API, because the thing being judged is the rendering itself.

Give the agent a socket whenever the software exposes one, even a bad one, because a flaky API still beats a reliable screenshot. When no socket exists but the target is a web app, go to the browser's protocol layer before you go to the pixels, and drive the DOM or CDP instead. Treat computer use as a bridge, never a destination. The moment a protocol ships, the pixel path becomes dead weight you are still paying for at five times the token rate.

Then hold the boundary condition honestly. The Astra run worked because it had low-level access most agents are never granted. It ran on a private server with the client source and the database in reach. Screenless control needs scoped protocol access, and that is an engineering grant you design and revoke, not a capability the model earns on its own.

Which makes the move to socket control a security architecture decision as much as an inference optimization, and it deserves to be built like one. Give each agent its own credential instead of sharing a human's. Scope that credential to the narrowest set of permissions the task needs, read-only wherever read-only will do, and a specific set of endpoints rather than a wildcard on the surface. Put it in a sandbox with a rate limit and a spend cap so a looping agent hits a wall instead of your invoice. Log every call with the agent's identity attached, so the audit trail I described earlier is actually attributable to something. And make revocation a one-click operation, because the day you need to cut an agent off is not the day to go looking for which key it is using. None of that is exotic. It is the difference between a socket you control and a socket that controls you.

WHAT TO DO TODAY

  1. Inventory every agent task that starts with a screenshot. For each one, find the endpoint, query, or file that produced the screen. Odds are it exists.
  2. Pick your most expensive pixel-driven workflow and prototype it against the raw interface for one week. Compare token spend and error rate. Keep both numbers, including the ones that embarrass the new approach.
  3. Separate orchestration from execution in at least one workflow. If your model is computing routes, coordinates, or lookups that a library or query could do deterministically, move that work out of the prompt and into code, then measure what the prompt is worth without it.
  4. Check whether your agent can verify its own actions. If its only evidence is screenshots, it is flying blind by design.
  5. Write down which of your systems have no protocol at all. That list is your real legacy debt, and it will be a bigger number than your model bill.
  6. If you are buying computer-use tooling, ask the vendor one question: what does this cost per action compared to an API call? Make them answer before you sign.

THE UNCOMFORTABLE QUESTION

Your agent has been reading screenshots like a human because you built it like a human. The demo that beat it ran blind, in 40 minutes, on packets and server files, and it finished the zone without dying once. How many of your agent tasks are you paying vision prices for because you never checked what was under the pixels?

Enjoyed this article?

☕ Buy Me a Coffee

Support PhantomByte and keep the content coming!

Build Real AI Infrastructure

PhantomByte teaches you to build real AI infrastructure yourself: local AI stacks, autonomous agents, multi-agent orchestration, web scraping, and custom tools. Step-by-step PDF tutorials you download, follow, and deploy. No subscriptions. No fluff. Just skills that ship.