TL;DR
89% of AI buyers exceed their initial budget, and only 10% of the time is it the vendor changing terms. The rest is usage compounding faster than anyone planned – which, in a usage or credit model, is the vendor's expansion revenue arriving as a surprise. The biggest overruns can't be forecast away, so the fix is control: let the customer see spend climbing and approve it before the bill lands. A hard cap and a soft cap stop the same dollars; only the soft one keeps the customer.
See the full budget and guardrail data in our 2026 AI Pricing report
A customer sets a budget for your AI product, and a few months later runs straight through it. They feel ambushed, and they start asking whether that is a reason to leave. For most AI software this is the normal outcome, and the cause is usually the customer's own usage rather than anything you did.
If you are the one pricing that product, it is still your problem: the customer feels the surprise, and the customer decides whether to renew. Handled well, the same overage reads as growth they chose. Controlling AI spend is a design problem – where, in your pricing, the customer gets to see spend climb and approve it before it runs.
Why do AI budgets go over?
Because customers adopt the product faster than anyone planned, and that compounding usage is rarely the vendor's fault. In Pricing I/O's 2026 survey of 296 software buyers, 89% had exceeded their initial AI budget – 45% significantly, and 44% moderately. At that rate an overrun is a property of the product rather than a string of mistakes.
The causes buyers report point inward, at their own usage outgrowing the plan they set at signing:
Read that table from the vendor's side and it inverts. Every leading cause – features driving activity, usage scaling, adoption spreading – is more consumption, and under a usage, credit, or consumption model more consumption is more revenue. To the vendor, the overrun is expansion, booked as growth. To the customer, it was a shock – and the customer's reaction is what the renewal turns on.
Why forecasting AI spend isn't enough
A forecast catches the small overruns and misses the ones that hurt. Buyers want predictability, and the obvious way to give it to them is a number at signing they can budget against. For the smaller overruns that holds. The large ones come from usage no plan could have held – the top causes in that table, features driving activity and usage scaling past plan – and no estimate catches those. Tightening the number just produces better-informed guesswork, because the usage was always going to outrun it.
So for AI spend, predictability rarely comes from a sharper forecast. It comes from control: the customer watches spend climb and approves the next increment before the meter runs, so the total never lands as a surprise.
A forecast is a guess made once, at signing. Control is a decision the customer keeps making as they go.
Should you cap AI usage?
Keep a ceiling, but make it one the customer can open. A hard cap and a soft cap stop the same spend; only one keeps the customer. Ranked by how many buyers want each control, the pattern is plain:
Buyers want to see the bill coming and decide for themselves. The dollar figure matters to them less than being caught off guard by it. A hard cap protects the budget by cutting the product off mid-task, often just as the customer is getting the most value from it. A soft cap protects it too, but by flagging the spend and asking before it continues, so the value and the revenue keep flowing by choice – an alert well before the ceiling, a one-click way to approve more, and a meter the customer can check on any day of the month.
A soft cap is only as good as its threshold. The approval step is friction sitting on the expansion you want to grow, so where you set it carries a revenue cost. Set it too low and the alert fires so often the customer learns to approve without reading – alert fatigue that protects nothing. Set it too high and the spend has already run before anyone can act. The threshold that holds is around 80% of the period budget, carried with a projection of where usage is heading, so the customer makes one informed call rather than fielding a stream of warnings.
The right guardrail depends on the cause and on the sales motion. Soft caps fit when usage is hard to forecast in advance, predictive alerts when adoption is spreading internally, throttling when the pricing unit itself is unclear. Motion matters as much. In a self-serve, product-led model no one is watching the account when usage spikes, so soft caps with one-click approval are the requirement and a hard cap is a liability – it breaks the product with no one there to recover the moment. In an enterprise, sales-led model an account team can step in, so a predictive alert that opens an expansion conversation does the work. What changes is who catches the overrun, the product or a person.
How do you build control into AI pricing?
Build spend control into the model from the start, rather than bolting it on after a customer complains. The vendors who win consumption pricing treat the guardrail as protection for their expansion revenue. A buyer who has been burned by an overrun, which is most of them, won't sign for an unbounded model. Give them a predictable base that covers normal use, metered units for whatever runs beyond it, an alert-and-approve step at the threshold, and a live meter that shows spend as it happens.
All of that manages the symptom. The deeper control is the pricing metric itself. A guardrail catches spend after it climbs; the value metric decides whether that climb tracks something the customer values in the first place. Tie the metric to a unit of value the customer feels – work completed, outcomes delivered, seats actively in use – and an overrun stops being a threat, because the customer who spent more got more. Tie it to an abstract credit or token the buyer cannot map to value, and no alert or cap can save it; the pricing only meters confusion. So build the guardrails, but get the metric right first. That order – a legible value metric, a predictable floor, then controlled expansion on top – is what AI buyers want most from pricing: the confidence to commit, and the room to grow.
Frequently asked questions
Why do AI budgets go over? Because usage outruns the plan. The top causes are AI features driving additional usage (67%) and usage scaling faster than expected (63%). Only 10% of overruns come from a vendor changing pricing or terms after the sale; the rest is the customer's own adoption compounding.
Can you forecast AI spend accurately? For moderate overruns, better forecasting and visibility help. For the largest overruns a sharper forecast changes little, because that usage was always going to outpace any number set at signing. Those cases call for real-time control: the customer seeing spend climb and approving it as it happens.
Should you cap AI usage? Keep a ceiling, but make it a soft cap the customer can approve past instead of a hard wall. A hard cap and a soft cap stop the same spend, but buyers rank hard caps (40%) and prepaid pools (28%) last and want soft caps with approval (62%) and predictive alerts (55%). A hard cap stops the vendor's revenue and the customer's at the moment the customer is getting the most value.
What is the difference between a hard cap and a soft cap in AI pricing? A hard cap stops the product once spend hits a set limit, cutting usage off mid-task. A soft cap alerts the buyer as they approach the limit and asks them to approve continuing, so the product keeps working and the spend stays a choice. Buyers prefer soft caps (62%) to hard caps (40%): they want to keep using the product without meeting a surprise bill.
How do you control AI spend? With guardrails that warn before they cut off, matched to why the budget broke: soft caps with approval, predictive alerts, throttling. Pair any usage pricing with a predictable base, a legible value metric, and real-time visibility, so the customer can always see the bill coming and approve the next increment before it reaches the invoice.



