Not All AI Agents Are Equal: A $172,000 Lesson in Denial Recovery Economics
By Vitali Khvatkov, Founder, Pinnacle Services Corporation · Study window: seven days in May 2026. Model versions and list prices are as of that date; the models are described but not named.
Cheaper per token is not cheaper per task. More capable in benchmarks is not more profitable in production. Over a seven-day window in May, we ran Recovr's denial-recovery agent through the same workload on three frontier large language models from two vendors — a mid-tier and a flagship model from the first vendor, referred to here as Model A and Model C, and the flagship model of a second vendor, Model B. Identical tools. Identical claims. Identical prompts. The cost spread was 5.2×.
That number matters more in revenue cycle management than almost anywhere else AI gets deployed.
The arithmetic of a five-cent margin
The math is unforgiving. An RCM firm keeps roughly five cents of every dollar it collects for a client. Every operating cost lands on that nickel, not on the dollar. The dominant cost-cutting playbook for the past decade has been offshore staffing — operators in India at roughly ten dollars an hour, working twenty to forty-five minutes per denial. That puts the human labor cost of recovering a single denial somewhere between $3.30 and $7.50.
AI agents do not have to match that number. They have to beat it — by enough to cover infrastructure, supervision, model overhead, and the margin compression that justified the build in the first place.
What three frontier models actually cost
Running Recovr's Claim Edit Agent — the subsystem that reads a denied claim, retrieves payer policy, identifies the root cause, drafts the correction, and submits it to the payer — across the three models, on production claims, with the same tool surface and the same prompts:
Model | Cost per claim recovery | Model rounds per recovery |
|---|---|---|
Model A (vendor 1, mid-tier) | $0.34 | 5.9 |
Model B (vendor 2, flagship) | $0.55 | 11.1 |
Model C (vendor 1, flagship) | $1.77 | 5.6 |

The naïve reading of vendor price cards predicts the opposite ranking. Model B's list input price is 2.4× cheaper than Model A's. The price card says Model B should win on cost. In production it lost by 60%. Model C is five times more expensive on paper than Model A — and on a structured, tool-using workload, it cost 5.2× more, almost exactly tracking the per-token premium. The premium did not buy a corresponding reduction in iterations. Model C converged in 5.6 model rounds; Model A in 5.9. The smarter model was not the faster model.
Where the gap comes from
Three forces explain the spread, and none of them appear on a marketing page.
The first is cache hit rate. The first vendor's context cache, properly engineered, served between 64% and 79% of input tokens at a deep discount. The second vendor's implicit cache served 24% to 29% at a shallower one. That single factor closed most of the headline price difference between Model B and Model A.
The second is iteration count. Each agent round re-sends the entire growing conversation as input. Model A completes the average denial recovery in 5.9 rounds; Model B takes 11.1. Twice the passes against a colder cache compounds — the lower per-token sticker price evaporates inside two extra trips through the prompt.
The third is tail risk. Models A and C clipped at thirty iterations even on the hardest claims. Model B produced a long-run rate above 10% and one runaway at eighty iterations, costing $3.50 on a single denial — ten times a healthy run. For an autonomous agent operating unattended, the worst-case run defines the cost story, not the median.
A fourth, invisible line item: the second vendor bills the model's internal "thinking" tokens as output at the output rate; the first does not. That alone adds roughly $0.05 per claim on Model B — about ten percent of the per-claim cost, and invisible to anyone reading only the published price card.
The line item that matters most is the one you wrote
Model A's $0.34 per claim is not a vendor's number. It is an engineering number.
Model A's baseline performance on this same workload — before we built the caching layer — was $0.55 to $0.65 per claim. At parity with Model B. The 60% improvement came from three disciplines: holding the static instructions and tool definitions byte-stable so the vendor's cache can reuse them, keeping the dynamic context blocks identical across iterations, and compressing resolved tool returns before they re-entered the prompt. The single largest economic move in the study was not model selection. It was context engineering. The right model run badly costs as much as the wrong model run well.
The point
At our projected production volume, the spread between the best and worst configuration is roughly $172,000 a year. That is meaningfully larger than the fully loaded salary of the engineer who eliminates it. It is also a noticeable fraction of the operating margin on a mid-sized RCM account.
The takeaway is not "use Model A." Vendors reprice, caching gaps close, and the rankings will change. The takeaway is that AI agents in revenue cycle management are not commodities, and treating them as one — picking the cheapest sticker price, or the highest benchmark score, and skipping the production engineering between that choice and a working economic unit — burns the five-cent margin the entire industry runs on.
Building a denial-recovery agent that competes with $10-an-hour manual labor requires deep expertise in both halves of the problem. Knowing CARC and RARC codes, MolDX policies, NCCI edits, and payer behavior is one half. Knowing how context caching, thinking-token billing, history compression, and iteration ceilings shape per-task economics is the other. Most teams have one half. A few have neither and don't know it. The ones that have both are the only ones whose AI investment in denial recovery will ever pay back the build.
Pinnacle Services Corporation builds Recovr, an AI denial-recovery platform for independent clinical laboratories, molecular diagnostics labs, and medical billing organizations. The full technical study — 138 production runs, 41.4 million tokens, complete telemetry — is available on request.


