$36.21 per run to $0.47. That is a browser agent's estimated model cost before and after, straight from a 144-run study Asana published with OpenAI in early October. The same agent went from more than 22.5 minutes per run to roughly four. That is 76 times cheaper and five times faster, and the model swap was not the lever.
Asana did not change the model to get there. Read that again. The fix was not a newer checkpoint, not a distilled small model, not a cheaper vendor. The team pointed a model at its own code, found that the harness was rewriting its own history at nearly every step, and stopped. They stopped invalidating their own cache.
Define the term in one line. A prompt cache means the provider stores the computed state of your request's opening bytes and reuses it the next time a request starts with the exact same bytes. Keep the opening bytes identical and most of the work is done before your request even arrives. Change them and the provider throws the cached state away and charges you full price for everything after the break.
The economics are why this matters more than any single harness bug. Every major provider that offers prompt caching behaves the same way at the boundary: an identical prefix is billed at a fraction of the fresh-token price, and a mutated prefix is billed at full price. OpenAI's own numbers on this study put the cached input at 5 percent of the uncached price. So a harness that rewrites its history every step is not paying a small inefficiency. It is paying full freight for a workload that could have been nearly free, and it is doing it on every call.
We priced the KV cache in "The Cache Is the Model: Why KV Cache Optimization Is the Most Underrated AI Infrastructure Play of 2026," where Inferoa AI measured a 97.8 percent cache hit rate on a real agentic coding task, and we showed how the cache dies between turns in "Your Agent's KV Cache Dies at Every Turn Boundary," where hit rates fell from 90 percent inside a turn to 55 percent across turn boundaries. Both notes explained the mechanism. Asana's study is the production receipt. Here is the leak, the fix, and the audit you can run tonight.
THE LEAK, IN PLAIN ENGLISH
Walk the failure step by step, because it is the same failure in almost every agent harness you will touch.
Asana's browser agent runs on StackAI, the platform Asana acquired, and it navigates websites, fills forms, and gathers information for customers. Every step, it collects page text and screenshots, and the history grows. To hold the request inside a 120,000 character budget, the agent dropped older screenshots and trimmed older page text. That is the reasonable thing to do when a context window looks like a scarce resource.
It is also the thing that was burning the money. Trimming edits an earlier part of the request. A request that resends the fixed instructions and tool definitions and then rewrites the history has changed its own prefix. A changed prefix cannot be reused from cache, so nearly every call recomputed the entire history from scratch, at full price.
The detail that makes it worse: the agent was already caching its fixed instructions and tool definitions. So the team was half right. The half they missed was the half that grows, which is the history of everything the agent reads.
The budget that was supposed to save money was the thing spending it.
WHY THE OBVIOUS FIX IS THE WRONG FIX
The intuitive move when an agent is expensive is to trim harder. Shrink the history, keep less, spend less. Asana's data says the opposite, and the number is not subtle.
They took the history budget from 120,000 characters to 480,000 characters, four times bigger, and the cost fell 76 times. On the same model, with the larger budget and the new caching policy, the per-run cost on GPT-6.1 Sol alone dropped four times, from $1.97 to $0.47. The constraint was never history size. It was history stability. Editing the middle of the transcript is the one thing a prefix cache cannot survive.
Sit with that, because it inverts the standard advice. The instinct to save money by keeping less context is exactly the instinct that breaks the cache and costs you more. Bigger, stable, and append-only beats smaller, thrifty, and constantly rewritten.
THE THREE MOVES
The fix is structural, and it is small enough to copy. Asana tested it across six caching and screenshot policies, each run three times on each of four models, inside the 144-run study. Three moves came out on top.

- Append-only history. Nothing already added gets edited mid-run. The transcript only grows at the end, so the prefix stays byte-stable and the cache keeps holding.
- Batch the pruning. Instead of cutting screenshots at every step, let them accumulate to 20 and then cut back to the most recent one. The prefix stays identical for long stretches between cleanups instead of changing on every turn.
- Size the budget for stability, not for thrift. At 480,000 characters, the agent kept enough history that it stopped re-reading pages it had already seen, and the larger budget is what let the caching policy pay off.
The result: 89 percent of the input came from cache at 5 percent of the uncached price, each call ran about three times cheaper, and every run in the optimized workflow completed the task with the correct answer. The fix made the agent cheaper and more accurate in the same move, which kills the standard excuse that you have to trade quality against cost.
The accuracy jump is the part to underline. On the smaller history budget, only three of 18 runs produced an answer at all. On the larger budget, all 18 completed, each with the correct answer. The agent was not just failing cheap. It was failing because it had been forced to forget.
THE STABLE PREFIX AUDIT
Here is the framework to keep, and it works on any agent harness against any provider. Three checks. Run them in order.
- The byte audit. Log the first N bytes of every request your harness sends. If the opening bytes differ from call to call, you have a leak, no matter what the provider's caching docs promise. The cache either matches your prefix exactly or it does not match at all. There is no partial credit.
- The mutation map. List every place in your code where the harness edits, prunes, compresses, or rewrites anything already in history. Every one of those is a cache killer. A context compressor that runs on the middle of the transcript is a cache killer wearing a cost-saving badge.
- The hit-rate meter. Read the provider's cache hit metric before and after any change. Asana's reached 89 percent. Under 50 percent on a long-running agent means the leak is you, and you are paying the tax.
WHAT TO DO TODAY
Pull your last 20 agent runs and diff the opening bytes of each request. That is the whole first audit, and it takes an afternoon.
Find every mid-run history edit in your harness and mark it append-only or batched. If you cannot make it append-only, make it happen on a schedule instead of on every step.
Read your provider's cache hit rate. If it is under 50 percent on a long-running agent, you are paying the 76x tax Asana escaped.
Note the side effect Asana reported, because it should change how you think about cost work. Fixing the cache fixed correctness too. Every run in the optimized workflow completed with the correct answer. Cheap and correct turned out to be the same fix.
One more thing worth stealing. Asana ran the investigation by pointing GPT-6 Astra in Codex at the problem and letting it run experiments against the codebase. Work the team estimates at one to two months by hand took about a week, and Asana's CTO described setting a goal before bed and reviewing the results in the morning. The lesson is not that you need that exact model. It is that the thing reading your agent's invoice should itself be an agent, and the same fleet you are trying to make cheap can be the fleet that finds the leak.
THE UNCOMFORTABLE QUESTION
Asana found a 76 times leak by pointing a model at its own logs for a week. How long has your agent's invoice been telling you the same thing, and who is reading it?
Get More Articles Like This
Getting your AI agent setup right is just the start. I'm documenting every mistake, fix, and lesson learned as I build PhantomByte.
Subscribe to receive updates when we publish new content. No spam, just real lessons from the trenches.
Build Real AI Infrastructure
PhantomByte teaches you to build real AI infrastructure yourself: local AI stacks, autonomous agents, multi-agent orchestration, web scraping, and custom tools. Step-by-step PDF tutorials you download, follow, and deploy. No subscriptions. No fluff. Just skills that ship.
