Microsoft and Hugging Face just ran 121,680 valid agent trials across 12 models. Of those, 79,853 failed the executable checks on the database state they left behind. Here is the part that should scare you. Of those failures, 67.24 percent ended clean. No error. A state-changing tool call. A confident "done."
The agent said it was done. The database disagreed.
We covered the diagnosis of this failure mode in "Read-Back Layer: Your Agent's Report Is Not Evidence," published October 3. ThinkingBox is the hard numbers behind that diagnosis, and the numbers say something worse than the diagnosis did. Even when the report reads fine, the reliability does not.
THE BENCHMARK THAT GRADES THE DATABASE, NOT THE ANSWER
ThinkingBox is a sandbox and benchmark that grades agents on the terminal backend state and side effects they leave behind, instead of the sentences they generate. The environment holds 507 stateful business workflows across retail, auto insurance, travel, neobank, consulting, and IT service management. Every workflow runs 20 times from an identical clean backend, and scoring is done by executing checks against the database, not by reading the agent's final message.
Two definitions carry the whole argument. pass@1 is the share of all attempts that succeeded, so it answers how the model usually does. pass@20 is the share of tasks solved at least once in 20 tries, so it answers whether the model can ever do the work at all. The benchmark also reports the literal observed 20 out of 20 count, which answers whether the model can always be correct. No estimator. No smoothing.
The motivating example is worth memorizing. A customer's $745 appliance sat stuck in a courier exception, fifteen days past its estimated delivery date. The agent did careful work. Nine tool calls. It pulled the order, checked tracking, looked up the customer profile, searched the refund policy twice, confirmed no ticket existed, opened one, documented the timeline, and read the policy correctly. Then it closed the ticket as resolved and asked whether there was anything else it could help with.
Two things were wrong. The carrier exception was still open, so the required end state was on hold, pending resolution. And the customer never got a real answer to what she actually asked. An AI grader checking tool calls would have seen nine well-formed ones. The database is what disagreed. The executable check that fails is a single field: the ticket status is solved where the required end state is hold.
THE 67 PERCENT YOU NEVER SEE

Walk the failure anatomy, because it is the whole reason this benchmark exists. Of the 79,853 failed attempts, 67.24 percent terminated cleanly, invoked a state-changing tool, and reported no final tool error. Nothing in the transcript flags a problem. When the executable checks inspected those silent failures, they found wrong field values in 77.61 percent of them, unintended extra effects in 43.30 percent, and missing required effects in 25.36 percent. Those findings overlap, because a single broken run can carry more than one defect.
For operators, silent failure is the primary baseline, not the edge case. If your monitoring only watches for thrown errors, failed tool calls, and non-empty error strings, you are blind to most of your agent's real failure rate. The agent that writes the wrong row and then tells you it finished is invisible to every check that reads the agent's own output.
A trajectory is a claim. Database state is the evidence. Repetition is the trust test.
PASS@1 IS A DEMO, PASS@20 IS A PRODUCT
Now the reliability spread, which is the citable finding of the whole campaign. Only three models hold on to most of their single-attempt score across the repeats. GPT-6 Astra retains 78 percent of its pass@1 rate, and Claude Opus 5.5 and Claude Opus 5 each retain 71 percent. At the other end, GLM-5.1, Kimi-K2.6, and DeepSeek-V4-Pro each keep about 8 percent. Same leaderboard, same first impression, and an order of magnitude between what they do once and what they do every time.
The breadth-and-consistency split is where the model choice gets brutal. Kimi-K3 has the broadest coverage of anything tested. It solves 93.89 percent of the benchmark at least once, which is 476 of 507 tasks, and only 31 tasks defeat it entirely. It is also among the least consistent, because just 68 of 507 tasks, 13.41 percent, succeed in all 20 attempts. Claude Opus 5 inverts the trade. It solves fewer tasks at least once, 79.09 percent, with 106 tasks defeating it entirely, but it completes 47.53 percent of the benchmark on every single attempt.
Put those two side by side and the buying lesson writes itself. Kimi-K3 solves 75 more tasks at least once than Opus 5. Opus 5 solves 173 more tasks consistently than Kimi-K3. Two models can look equal in a demo and differ by an order of magnitude in production. Never accept a single-run demo as evidence. Demand the repeat count.
The benchmark even kills the reflex to buy the newer model. Claude Opus 5.5 scores higher than Claude Opus 5 on the every-attempt average, 67.16 percent against 66.50 percent, and solves more tasks at least once. It passes exactly the same number of tasks on all 20 attempts, 241. Half a point of headline accuracy bought no additional dependability at all.
THE STATE-CHECK RULE
Here is the citable framework, and it fits in one sentence. An agent task is done when the system of record says done, and it is reliable only at the run count you actually measured.
Two halves, both required.
First, terminal-state checks. Define the required end state as executable assertions against the database, not as a human reading the agent's summary. Check all three failure modes, because those are exactly where the failures live: wrong values, extra effects, and missing effects. The benchmark accepts any trajectory that produces the right outcome and rejects wrong, missing, or extra effects, which is the correct design. You do not care how the agent got there. You care what the database holds when it stops.
Second, repeat discipline. Run every workflow N times from a clean state and track the literal N-for-N count, not the pass rate on one run. A workflow is production-ready at the N you will actually run it, not at N=1. For most production work that number is far above one, and the gap between pass@1 and 20 out of 20 is where your incidents live.
Third, the implementation pattern, because a mandate without a shape is just an opinion. Run each episode against an ephemeral test container or a shadow copy of the database, never the live record. Snapshot the schema and rows to JSON before execution, snapshot again after, and diff the two. Assert that the diff matches the expected effect set exactly. That snapshot-diff loop is the mechanism behind every executable check, and it is cheap enough to run 20 times.
Separate the two kinds of assertion, because they fail differently. A state invariant is a predicate that must hold in the final state no matter which trajectory the agent took, such as ticket.status == 'hold'. A side-effect payload check verifies the exact values a specific action was supposed to write, such as the refund amount or the carrier exception flag. Invariants catch missing and wrong outcomes. Payload checks catch the subtle value errors that 77.61 percent of silent failures carry. Test both, because the first tells you the workflow ended wrong and the second tells you how.
Then handle partial state honestly, because silent failures do not always leave a clean record. The 43.30 percent extra-effect rate is half-mutated state: a row updated, a second row created, and no error anywhere. Wrap each workflow's writes in a transaction with a rollback hook that fires when the terminal assertion fails, so a failed episode cannot contaminate the next one. Make the state-changing tools idempotent, keyed so a retry after a rollback does not double-apply. Without rollback and idempotency, your 20 runs are not 20 identical clean runs, and the reliability number you measured is not the one you think you have.
CLOSING THE LOOP WITH THE FAILURES
ServiceNow's AutoSynthData, published the same week, closes the loop. It turns a target model's failures into training data instead of scraping more generic instruction sets. The pipeline runs a weaker target model and a stronger teacher model on the same diagnostic tasks, looks for where the target fails and the teacher succeeds, and generates new validated tasks that exercise exactly that capability. As the model improves, the curriculum shifts toward whatever it still gets wrong.
The numbers are real and they are modest, which is the honest framing. On the EnterpriseOps Gym Hybrid domain, AutoSynthData generated 2,000 synthetic training samples in about 18 hours. Fine-tuning Gemma-4-26B-A4B-it on that data raised mean pass@1 by 7.2 percentage points, a 35 percent relative gain, and lifted verifier success from 63.01 percent to 68.55 percent. That closed 59 percent of the original gap to the reference model. Note the scope: EnterpriseOps Gym is a synthetic enterprise sandbox, so those are gains inside a controlled environment, not a live-enterprise trial.
The pipeline lesson for teams without a training budget is the part nobody says out loud. Your executable checks are not just a gate. They are a dataset. Log every silent failure with its expected state and you have the raw material for fine-tuning, or for prompt regression tests you can rerun on every change. The environment and the data behind ThinkingBox are public with pinned hashes, so a canonical result can be reproduced rather than asserted. Reproduce, do not assert.
WHAT TO DO TODAY
- Pick one agent workflow you own and write the terminal state as an executable check against the actual database, not a summary review.
- Run it 20 times from a clean backend before you call it reliable. Record the 20 out of 20 count, not the pass rate on one run.
- Audit your monitoring for silent failures. If it only alerts on errors, it is blind to the 67.24 percent case.
- When a vendor demos an agent, ask one question. What is the pass@20? If they only have pass@1, you are looking at a demo, not a product.
- Start a failure log keyed to expected database state, so every silent failure becomes training or regression material.
THE UNCOMFORTABLE QUESTION
Your agent succeeded the one time you tested it. If two thirds of agent failures end with a clean "done" and no error, how many of your databases are already wrong?
Get More Articles Like This
Getting your AI agent setup right is just the start. I'm documenting every mistake, fix, and lesson learned as I build PhantomByte.
Subscribe to receive updates when we publish new content. No spam, just real lessons from the trenches.
Build Real AI Infrastructure
PhantomByte teaches you to build real AI infrastructure yourself: local AI stacks, autonomous agents, multi-agent orchestration, web scraping, and custom tools. Step-by-step PDF tutorials you download, follow, and deploy. No subscriptions. No fluff. Just skills that ship.
