Nvidia published a paper on a system called SoL-Pi, and the headline number is the part that should stop you. Token usage drops by almost half while performance stays roughly the same. That is 50 percent savings against Codex. It is 54.3 percent against Claude Code on EdgeBench.
Here is the uncomfortable detail. SoL-Pi does not touch the model. It does not quantize anything. It does not swap in a cheaper checkpoint. It does not tune a single attention kernel. It optimizes the agent harness, the control layer between the agent and its environment, and that layer was quietly eating half your budget.
We priced this in Note #195: your harness is paying for components it does not need. We proved the harness moves your benchmarks 4.3 times more than your training recipe in Note #184. Nvidia just automated the repair.
THE HARNESS IS THE PART YOU NEVER INSTRUMENT
Define the term once, then hold onto it. The harness is the code that decides how your agent sees state, calls tools, assembles context, verifies results, and aborts bad runs. Everything the model does passes through it. Nothing the model does bypasses it.
Almost nobody profiles that code. For the last two years the efficiency work went into the model side: faster attention kernels, quantization, cheaper model substitution, better serving infrastructure. Every one of those attacks cost per token. SoL-Pi attacks tokens per task. Those are two different ledgers, and only one of them shows up on your invoice.
You can see the split in the numbers. On EdgeBench's 51 public tasks, the most efficient SoL-Pi variant uses 49 percent fewer tokens and still reaches 93.7 percent of the original Pi harness score, landing at $894 instead of $1,339. The variant that prioritizes raw performance beats Pi's score by 5.3 percent and still saves tokens. Across both, token traffic falls by 44.7 to 49 percent. The authors put the hourly savings at $8.75 to $13.50 against native Codex and Claude Code harnesses, and $4.36 to $5.71 against Pi, at current API prices.
Reconcile one number before you quote it, because a staff engineer will spot the gap. Token traffic falls by 49 percent, yet the bill drops by about a third. That is not an accounting trick. The complete stack cuts cache-read traffic from 2.1326 billion tokens to 1.0605 billion, but cache-write traffic rises from 0.0141 billion to 0.0316 billion, and cached input still costs money. Shorter context forces prompt-cache rewrites that get paid for on a later turn. Input and output tokens also carry different prices, and the mechanisms compress one harder than the other. The savings are real. They are just not proportional to the token count.
Same model. Same tasks. Same weights. Different control layer. That is the entire trick.
WHY HARNESS OPTIMIZATION IS HARD: THE COUPLING PROBLEM
Tool usage, context management, verification, and abort logic are tightly coupled. A change that saves tokens in one place can trigger errors somewhere else, or it can simply push the cost into a later phase of the run. You trim context at step three, the agent loses the detail it needed at step nine, it re-reads the file, and you paid twice for the same information.
That is why hand-tuning a harness stalls. Every local fix risks a global regression, and you cannot see the regression until the run finishes. Humans end up sifting through thousands of execution traces and translating recurring failure patterns into code by hand. It is slow, it is boring, and it is exactly the kind of work a system should be doing to itself.
The coupling is not theoretical, and you do not have to take Nvidia's word for it. In August, the tooling company Composio ran Deepseek V4 Flash across four different agent frameworks, including Claude Code and the Pi-based Oh My Pi. The same model solved the same tasks at a cost that varied by nearly 3x depending on which harness wrapped it. Nothing changed except the control layer. Composio proved the problem existed across frameworks. SoL-Pi provides the automated loop to solve it.
The EdgeBench numbers are the proof one level down. Change one part of the control logic and the whole run moves on you.
HOW SOL-PI WORKS: THE HARNESS AUDIT LOOP
SoL-Pi puts a research AI on the problem. That agent watches another agent's execution traces, proposes changes to the harness, and tests each proposal in a prepared environment. Capability checks and efficiency checks decide which candidates survive. Nothing survives on a hunch.

The scale is the interesting part. Across 535 executable environments, the system explored 152 directions, including 495 tasks derived from GitHub issue and pull request pairs plus 40 synthetic test cases. That produced more than 3,000 runs and over 60,000 agent to environment interactions. Four mechanisms came out the other side.
- Action Fusion merges two consecutive steps into one call, which kills an entire model call when a code edit is immediately followed by a test run.
- Online Context Compact trims accumulated context after each planning step whenever it can do so without losing something important.
- ObservationPack archives long tool outputs and leaves a short summary behind instead of resending the full text every step.
- The Evidence-Preserving Reducer routes large error and test logs to a cheaper model, boils them down to the findings that matter, and runs an automatic verification pass to catch the clue that slipped through. That pass is deterministic and it checks the receipt's schema, source hash, exit status, exact quotes, and size. If verification fails, credentials look suspect, or the receipt saves no space, the harness falls back to the raw log. It trades cost for precision, never silence.
Those mechanisms are the output. The process is the thing you should steal, and it deserves a name. Call it The Harness Audit Loop. Trace, propose, measure, keep or kill. It is the same keep-what-survives pattern as a CI gate: nothing merges on vibes, everything merges on measured performance held at lower cost, and the kill rate is the point.
One design detail matters more than any single mechanism. SoL-Pi separates search feedback from evaluation completely. The harness was frozen before the system ever touched the held-out tasks, and those results never fed back into the search. That wall matters because earlier work showed that automatically optimized harnesses tend to overfit to the tasks they were built on and deliver almost nothing on unfamiliar ones. If you skip the wall between search and evaluation, you are not optimizing a harness. You are memorizing a benchmark.
There are honest trade-offs here, so read them before you quote the headline. On 63 CPU tasks from Terminal-Bench 4, SoL-Pi solved 15 while Codex and the original Pi each solved 18, even though total cost still came in about a quarter lower. The system was built on GPT-5.6 Sol trajectories and then applied to Opus 5 without changes, where it held 94.3 percent of Pi's performance, but the mechanisms fired less often, because the harness had been tuned against one model's behavior. Shorter context can also cut prompt cache reuse, which claws back some of the gain. This is not a free lunch. It is a real one with a receipt.
WHAT THIS MEANS FOR YOUR BILL
Do the math in plain sentences. If your coding agent spends $10,000 a month in tokens and half of it is harness overhead, your model is not the cost problem. Your control layer is. You can negotiate the best per token price in the industry and still burn the same money, because the waste is happening before the meter you keep staring at.
Buying a cheaper model on top of a wasteful harness is paying twice for the same mistake. You cut the price of every token and you keep requesting twice as many as you need. Meanwhile the pressure keeps building, because agentic token usage has grown 14x since February and nearly 70 percent of that traffic is cached prompts, which means the compression gains and the cache discounts are fighting over the same dollars.
The cheapest token is the one your harness never requests.
WHAT TO DO TODAY
- Pull the traces. Before you touch the model, capture your agent's decision traces for a representative week of runs. You cannot audit what you never recorded.
- Count the phases. Map where tokens actually go: state assembly, tool calls, verification, retries, aborts. Find the biggest block and start there.
- Change one coupling at a time. Never bundle harness edits. A bundle makes a regression untraceable, and you will not know which change broke the run.
- Gate on performance, not vibes. A harness change ships only if cost falls and measured task performance holds. No exceptions, no "it feels faster in testing."
- Run a keep-or-kill loop. Propose changes the way SoL-Pi does: measure, keep the survivors, kill everything else. If nothing gets killed, you are not searching, you are shopping.
- Re-audit after every model swap. A new model changes the optimal harness, because context windows resize, sensitivity to prompt structure shifts, and reasoning depth per step moves. Nvidia watched it happen: the same harness held 94.3 percent of Pi's performance on Opus 5 but triggered its mechanisms less often, because it had been tuned on GPT-5.6 Sol trajectories. The control logic that fit one model does not fit its replacement. Last month's tuning is this month's waste.
THE UNCOMFORTABLE QUESTION
Nvidia cut agent token spend in half without touching a single weight. So what exactly has your team been doing at the model layer this whole year, solving the cost problem or avoiding it?
Get More Articles Like This
Getting your AI agent setup right is just the start. I'm documenting every mistake, fix, and lesson learned as I build PhantomByte.
Subscribe to receive updates when we publish new content. No spam, just real lessons from the trenches.
Build Real AI Infrastructure
PhantomByte teaches you to build real AI infrastructure yourself: local AI stacks, autonomous agents, multi-agent orchestration, web scraping, and custom tools. Step-by-step PDF tutorials you download, follow, and deploy. No subscriptions. No fluff. Just skills that ship.
