Generative AI

Why your agents are burning budget and not driving outcomes

Author
Tomer Benami
August 13, 2026

Last summer, Gartner predicted that more than 40% of agentic AI projects would be canceled by the end of 2027, blaming escalating costs, unclear business value, and inadequate risk controls. A year on, that number keeps resurfacing in board decks and budget reviews, usually stripped of its date and treated as fresh news. The prediction has aged well enough that it hardly needs the update.

If you approve or defend an agent budget, the reflex when you read that is to worry about the technology. The models aren't ready yet. The vendor oversold. Wait for the next release and the returns will show up.

I'd push back on that. Some projects do fail on capability: the task was genuinely beyond the agent, and no amount of measurement would have saved it. But that's not what's driving a 40% cancellation rate. The models are good enough to be useful right now, and they are improving faster than most budgets can absorb. For most of the programs getting cut, the agentic AI ROI problem in front of you is less a capability problem than a measurement one. They aren't dying because the agent couldn't do the work. They're dying because nobody could prove whether it did.

That distinction decides which side of the 40% your program lands on. So it's worth being precise about what your team is actually measuring today, and what they should be.

You're measuring activity, not outcomes

Look at the dashboard your agent team reports against. It almost certainly tracks tokens consumed, tasks attempted, average latency, requests served, and cost per run. All real numbers. All easy to pull. And not one of them tells you whether the agent accomplished the thing you bought it to accomplish.

That is the gap between activity and outcomes. Activity is what the agent did. Outcomes are what changed in the business because it did. An agent can burn through a fortune in tokens, complete every task it was handed, and still resolve the ticket incorrectly, misfile the contract, or route the customer to the wrong place. The activity metrics stay green the entire time.

Think about how you'd judge a sales team. You don't evaluate a rep on calls dialed and emails sent. You evaluate them on deals closed. Dialing the phone all day is motion; closing revenue is the result you're paying for. Judging an agent on tokens burned and tasks attempted is judging the rep on dial count — measuring the effort instead of the result, and hoping the two correlate.

Sometimes they do. Often they don't, and the gap is exactly where budget quietly leaks: real money spent on activity you can see, while the outcomes you're actually buying go unwatched.

ROI has a denominator and a numerator, and every program instruments the denominator well: tokens and inference, infrastructure, the build, the human oversight the agent still requires, the hours spent cleaning up its mistakes. You can produce that figure to the dollar. The numerator is benefit: labor displaced, cycle time compressed, revenue influenced, reversals avoided. That's the half almost nobody instruments. So programs end up with a precise, growing cost figure divided by a benefit that's mostly asserted. When someone at the annual review asks what this agent returned, "it processed four million tokens last quarter" is a denominator masquerading as an answer, and it does not survive the room.

This is what "unclear business value" looks like from the inside. Not an agent that failed, but an agent whose success was never measured in terms anyone upstairs recognizes as value. And it's widespread: in Foundry's 2026 State of the CIO survey of 662 IT leaders, fewer than one in five said their AI initiatives had met or exceeded business goals. The barriers they named most were ill-defined ROI metrics and murky strategy, not the underlying technology.

Why outcomes go unmeasured

If outcome measurement is so obviously the thing that matters, why does nearly everyone default to activity instead?

Because activity is easy to instrument and outcomes are hard. Tokens, latency, and task counts stream out of the system automatically. Your platform emits them whether or not you ask. Standing up that dashboard is close to free.

An outcome is a different kind of thing. To measure one, you have to answer a genuinely hard question: did the right thing happen? Did the refund process correctly? Did the agent's decision hold up against what a correct decision would have been? Nothing streams that out of the box. Somebody has to define what "right" means for each task and then check the agent against it.

So the easy measurement crowds out the important one. Teams instrument what the tooling hands them and report the numbers they have instead of the numbers that matter. It isn't negligence; it's the path of least resistance, taken one reasonable decision at a time, until you've built a reporting stack that describes everything about your agent except whether it worked. And value you can't measure is value you can't defend when budgets tighten.

Where production measurement runs out

To know whether an agent achieved an outcome, you need something to check its work against: a known-correct answer, the equivalent of an answer key. The obvious question is where that key comes from, and the obvious answer starts with an objection to my own argument. In production you measure realized outcomes all the time without any answer key. Resolution rate, cost per resolved ticket, escalation rate, reversal rate, CSAT: these are real outcome metrics, and if you're running an agent at scale you're almost certainly watching them. So it would be wrong to say you can't measure agent outcomes without controlled data. You can.

But production measurement has two limits that matter to anyone signing off on the spend. The first is timing: those numbers only exist after the agent is live and handling real customers and real money, so the first rigorous read on whether it works arrives after you're fully exposed to the cost of it not working. The second, and larger, is attribution. When resolution rate ticks up, was it the agent, the new routing rules, a seasonal dip in hard tickets, or the two experienced reps you added that quarter? Production has too many moving parts to cleanly isolate the agent's contribution. You can see the aggregate moved; you often can't prove the agent moved it.

A controlled answer key buys you exactly what production can't. It lets you measure the outcome before you're exposed, and isolate the agent's effect cleanly, because the environment holds everything else constant. You defined the correct outcomes when you built the dataset, so a passing run is the agent's doing and a failing run is too.

This is the strategic question worth carrying into your next planning conversation: do we own the ground truth we'd need to prove this agent works?

For most teams the honest answer is no. They test against production data or borrowed datasets because that's what's available, but production data is messy and partial: it's the reality the agent has to handle, not a standard you can grade against, because no one holds a clean account of the correct outcomes buried inside it. Testing against it tells you the agent ran, not that it was right.

Proving an agent works before you stake the business on it has a prerequisite most programs never satisfy: a verifiable source of truth, which means data you control rather than data you borrowed. The teams that get there stop asking only whether their model is good enough and start asking whether they own the data they'd need to prove it. It reframes part of the budget conversation from a debate about model quality, which you can't do much about, into a question about your own data foundation, which you can.

It's also part of why so many well-run programs still can't produce a clean ROI number early enough to defend the spend. They invested in the model and the orchestration and skipped the layer underneath: the controlled data that lets them prove the agent's contribution before production does it for them, expensively. Escalating cost with value you can't isolate until late is precisely the profile of a project heading for cancellation.

The question to bring to your next planning meeting

None of this requires a better model or a bigger budget. It requires deciding, before the next tranche of spend goes out, whether you can answer one question about the program: do you own the ground truth you'd need to prove this agent works, on your terms and your timeline, rather than waiting for production to render its verdict after you're already committed?

If the answer is yes, you can put a defensible number next to the cost you're already tracking, and the ROI conversation becomes arithmetic instead of argument. If it's no, that's the gap to close first, ahead of further investment in the model or the orchestration around it. The data foundation is the part you actually control, and the part that decides whether you can prove value before the budget review does it for you.

That's the line between the programs that survive the 40% and the ones that don't. The teams treating their data foundation as the thing to get right, rather than the thing to figure out later, are the ones who'll still have a program to defend in 2027. Building that foundation from controlled data is the problem Tonic Fabricate exists to solve, through synthetic environments an agent can be tested in against known outcomes; the engineering side of the argument is worth your team's time.

Tomer Benami
CFO & VP of Business Development

Tomer Benami is the VP of Finance and Bizops at Tonic.ai where he brings a blend of core finance expertise, operational savvy, and vision to go-to-market activities. With a proven track record of serving as the senior-most finance leader at companies such as VirtualHealth and Apploi, Tomer enjoys partnering with executive teams, steering organizations towards strategic goals and delivering meaningful results. Beginning his career at KPMG and holding a Master's Degree from the University of Washington, Foster School of Business, he is enthusiastic about the transformative potential of AI while advocating for its responsible and ethical utilization in shaping our future.