Discussions about AI pricing usually start with the billing model. Seats, usage, credits, workflows, outcomes, or some hybrid combination of the lot.
Fair enough. But I think that starts one step too late.
The harder question is how to set a stable, commercially meaningful price when the cost of delivering the product is still moving underneath you. Buyers want a result they understand and enough predictability to budget for it. Vendors, meanwhile, are dealing with model costs, workloads and architectures that can change materially over the life of the price they have just promised.
I’ve seen this first-hand in conversations with AI founders. One had built the first version of his AI-native application using Claude Sonnet. At one point, it was costing him roughly US$40 per user per week while he was charging around US$38 per month. The maths was not especially subtle.
Once he had enough traction, he rebuilt the architecture around a much cheaper open-weight model. His inference cost later fell to roughly US$3.70 per active user, with broader server costs also falling significantly. The product had not suddenly become ten times more valuable. The economics of producing it had changed.
His dilemma was simple enough to describe and much harder to solve. His cost structure was changing rapidly, yet he still had to form a view on what it might look like in the future and set a price customers could live with.
Another founder I have been working with runs a voice AI platform. His situation is different but lands in much the same place. Today, much of the workload runs through external infrastructure and APIs, partly because suitable GPU capacity in Australia is constrained and data residency matters. At greater scale, more of that workload may move onto self-hosted infrastructure. If it does, the cost structure shifts from predominantly usage-driven external charges towards a mix of fixed capacity and lower incremental processing costs.
Neither founder can wait for the technology stack to settle before setting a price.
The customer would like to know now.
The problem is not simply that AI costs more
It is tempting to reduce this to inference. Traditional SaaS enjoyed forgiving marginal economics; AI introduces significant compute costs; therefore gross margins come down. True enough. But that is the easy part.
AI cost-to-serve is made up of several costs that behave differently. Take a voice AI product handling calls for a medical practice. One call might create telephony charges, speech-to-text costs, model inference, calls into a practice management system, text-to-speech and underlying infrastructure expense. If something goes wrong, perhaps a human gets involved as well.
The customer experiences one call. The vendor may have paid for half a dozen things to deliver it.
Those costs do not all move in sync. Speech processing may broadly follow minutes. Inference depends on model choice, tokens and context. External APIs may charge per transaction. Human intervention appears only in certain cases. GPU infrastructure behaves more like capacity: relatively fixed until the limit is reached, then another block of cost arrives.
The workload itself also varies. A routine appointment booking may require very little work. Another call may need more context, several tool calls, multiple retrievals and the occasional retry before anything useful happens. Both can still be recorded as one successfully handled call.
Sophisticated founders already understand this. They do not stop at the average. They look at P90, P95 and sometimes further into the tail because that is where the expensive workloads tend to live. P90, for example, is the cost level below which 90% of workloads fall, leaving the most expensive 10% above it.
The more interesting problem is that today’s P90 is not necessarily next year’s P90.
A model changes. Routing improves. Context windows get longer. A product release introduces another external tool. Customers discover new use cases. A cheaper model makes it economical to use the product much more heavily.
You might route simple workloads onto a cheaper model and drive median cost down sharply, while the product simultaneously becomes capable of handling more complex work at the top end. P50 improves. P90 moves the other way.
Or both improve. Or neither.
The point is not that AI costs are unknowable. They are measurable, and they should be measured properly. The problem is that the thing you are measuring has a habit of changing.
The price, unfortunately, tends to be rather more stationary.
The buyer does not care about your cost stack
While the vendor is thinking about model calls, context windows, APIs and infrastructure, the buyer is looking at the product from a different angle. They want a result.
A medical practice may care about calls answered, appointments booked and revenue recovered. A customer support team cares about issues resolved. A legal team may care about matters reviewed. A finance team may care about invoices processed or reconciliations completed.
They do not care which model produced the result, and there is no particular reason they should.
They also want to know what the product is likely to cost. Recent Benchmarkit/Pricing I/O buyer research reinforces that point: predictability of total cost matters greatly to buyers, often more than simply securing the lowest headline price.
That creates the tension at the centre of AI pricing. Internally, the cost base argues for flexibility. Externally, the customer wants a result they understand and a bill they can predict.
Passing every bit of volatility straight through to the customer would certainly make the vendor’s life easier. It may make the product less pleasant to buy.
This is why outcome pricing is attractive. Intercom’s Fin AI Agent is a useful example. A standard Fin resolution is priced at US$0.99. That is a much more meaningful commercial unit than asking a customer to buy tokens or model calls. The customer understands what a resolution is and, with some idea of expected volume, can form a reasonable view of spend.
It is neat.
But the neatness is on the customer side.
Underneath one resolution might sit a short retrieval and a couple of model calls. Another may involve a longer conversation, more context, several retrievals and multiple tool calls. The customer still sees one resolution. Intercom gets everything underneath it.
That does not make outcome pricing wrong. Quite the opposite. It may be exactly the right commercial metric. But moving pricing closer to customer value does not remove the cost problem. It transfers more of that complexity to the vendor.
This is where I think much of the AI pricing discussion stops too early. There is plenty of attention on what the customer values and what they are willing to pay. There is less attention on how the cost of producing that result is likely to behave over the period for which the vendor has committed to the price.
That distinction matters.
If you charge US$0.99 for a resolution, the issue is not merely whether today’s average cost leaves enough margin or whether today’s P90 looks comfortable. It is whether that price still works across a reasonable range of cost curves over the next year or two.
A pricing metric can be perfectly aligned with customer value and still be fragile underneath.
The bundle is part of the economics
This is why I prefer to think about pricing architecture rather than pricing model. The billing metric matters, but so do the bundle, included allowance, overages, minimum commitments, tiers, contract duration and the ability to reprice.
These may look like packaging decisions. They are really decisions about how much uncertainty the vendor is prepared to carry.
Consider an AI product charging $2,000 per month for up to 1,000 completed resolutions. The buyer gets a clear result at a predictable cost. For the vendor, the bundle sets the risk boundary.
Within those 1,000 resolutions, some jobs will be cheap and others expensive. Some customers will use almost the entire allowance; others will not. The bundle pools that variability. An overage beyond 1,000 protects against unexpected volume. Different tiers can separate materially different workload profiles. A fair-use provision can protect against the customer who discovers that one “resolution” can consume an heroic amount of compute.
Contract length matters as well. A twelve-month fixed price means living with your assumptions for twelve months. A twenty-four-month fixed price is a more confident statement about a cost curve that may not share your confidence.
Seen this way, “hybrid pricing” is not much of an answer by itself. A hybrid structure is only useful if you know what uncertainty it is supposed to contain. Volume risk? Workload complexity? Exposure to an expensive feature? Differences between customer segments? Or simply the amount of monthly variation the buyer is willing to tolerate?
The structure should follow the economics.
Not the fashion.
The real problem is duration
Even if you understand today’s economics extremely well, pricing decisions often outlive the cost assumptions used to justify them.
That, I think, is the part that deserves more attention.
You may know today’s average cost, P90 and expensive tail down to the decimal. Useful. But if you commit to a twelve-month price, you are also making a twelve-month judgement about model costs, architecture, workloads and customer behaviour.
Those assumptions do not have twelve-month contracts.
Suppose an AI company currently spends $20 per customer per month to deliver its product. One future is pleasant: model prices fall by 40%, customer behaviour stays broadly unchanged, and cost-to-serve falls to $12. Margin expands. Everyone looks clever.
Another future is less cooperative. Model prices still fall by 40%, but the product improves enough that customers use it twice as much. The unit cost of AI falls, yet total cost per customer rises to $24.
A third future looks different again. The company moves from external APIs to its own infrastructure. Incremental inference cost drops materially, but it is now paying for fixed GPU capacity whether utilisation is 40% or 90%.
Exactly the same customer price can sit on top of all three futures.
That is why setting an AI price today is partly a view on tomorrow’s economics. Not a precise forecast. That would be optimistic. But a range of plausible outcomes against which the price still needs to work.
The longer the commercial commitment, the more uncertainty the vendor absorbs. That does not mean repricing whenever a model provider changes its rates. Buyers quite reasonably expect more stability than that, and a pricing page that behaves like a commodities screen would become tiresome very quickly.
It means leaving enough room in the pricing architecture for the economics to move without breaking the commercial promise.
For me, that reduces the problem to three decisions.
The first is understanding the cost curve properly. Not just today’s average and not just today’s P90, but what drives the distribution and what could move it.
The second is defining the customer promise. What result are you charging for, and how much predictability are you prepared to provide around that result?
The third is setting the risk envelope. How much variation will you absorb inside the price, and where will bundles, tiers, overages, minimum commitments or contract terms begin transferring that risk back to the customer?
That is a different question from asking whether seats, usage or outcomes are best. There is no universally best model. The right architecture has to work for today’s economics, where you think they are heading, and how much uncertainty you are willing to carry.
AI buyers want something reasonable: a result they care about, at a price they can understand and predict. AI vendors have to give them that while the cost of producing the result is still changing.
The aim is not to find a price that perfectly fits today’s cost-to-serve. Today’s number may have a fairly short shelf life.
The job is to set a price customers can understand and plan around, while leaving enough room for the economics underneath it to change.
Because the price you commit to today may well outlive the cost curve you used to justify it.