For two decades, software has been the best-behaved line in a firm's budget. A licence had a price, a seat count, and a renewal date. Finance could forecast the cost a year out and be right, because the cost had almost nothing to do with how much anyone used the product. The analyst who lived in the tool and the partner who opened it twice a quarter cost the firm exactly the same.
AI spend does not behave like this, and the difference is structural, not teething trouble. Frontier models are metered. Every request consumes tokens, every token is billed, and the invoice at the end of the month is a precise record of how the firm actually worked. The cost no longer follows headcount. It follows behaviour: the analyst mid-diligence and the partner who reads the finished memo no longer cost the firm the same.
That would be manageable if behaviour were disciplined by price at the point of use, but it is not. To the person typing, any single request feels free, because nothing about the moment of use signals otherwise. So usage flows towards whatever is easiest rather than whatever is worth the most. Entire data rooms are pasted into a context window when two documents would have answered the question. The most capable model is applied to formatting work a far smaller one handles. Ten drafts are requested where one would have been read. None of these decisions is careless on its own; in aggregate, they are the budget.
Why budgets die in spring
The result is the pattern we built a practice around: a compute budget planned for the year that is substantially gone by the end of the first quarter. When it happens, firms tend to reach for one of two remedies, and both are wrong. Rationing access teaches the team that the tools are scarce, which quietly undoes the adoption the firm spent months building. Topping the budget up teaches the firm nothing at all, because next year's overrun is already in motion.
The real problem is neither the size of the budget nor the enthusiasm of the team. It is that nobody can say what the spend bought. An invoice denominated in tokens answers the question of how much with great precision while saying nothing about for what. No other line in the firm's accounts would survive that standard.
The wrong unit
Tokens are the correct unit for billing and the wrong unit for management. A partner has no intuition for what a million tokens ought to cost or produce, and no dashboard will build one, because the unit does not map to anything the firm makes decisions about. What a partner can price instantly is work: a memo reviewed, a data room summarised, a board pack drafted, a screen of comparables assembled. Those are the units the firm already thinks in, negotiates in, and staffs against.
A partner cannot price a token. Every partner can price a memo.
Workflow-level costing is the translation between the two. Instead of a single undifferentiated AI line, spend is attributed to named workflows, and each workflow acquires a unit cost: what first-pass diligence costs per target screened, what an adversarial review costs per investment committee memo, what a data room summary costs per room. Set beside the analyst-hours the work displaces or the speed it adds, each line supports an actual decision: keep, improve, or stop.
Attribution also changes what expensive means. A workflow with a high unit cost that saves a deal team two days per target is cheap. A workflow with a negligible unit cost that produces output nobody reads is pure waste, and there is more of the second kind than most firms expect, precisely because each instance is too small to notice. Without attribution, the worthless spend hides inside the aggregate, and the valuable spend gets caught in the same squeeze when the cuts eventually come.
Measure, attribute, prune
The discipline itself is not complicated, though the order matters. Measurement comes first: usage made legible by team, by workflow, and by model, rather than arriving as one number on a monthly invoice. Attribution comes second, and it is the step firms skip: every token belongs to a named workflow, so the question in the budget review is never why AI is so expensive, but whether the memo-review workflow is worth what it costs. Pruning comes last, and it applies to workflows, not people. Some work moves to smaller models that produce the same output for a fraction of the spend. Some prompts get shorter and some context gets tighter. Some workflows are simply stopped, because a line that earns nothing should not survive the reading of the P&L that finally made it visible.
Pruning cuts in both directions, and this is the part an austerity-minded reader will miss. A workflow that clearly earns its cost should be fed, not rationed. Our view is that the honest outcome of the exercise is rarely a smaller budget; it is the same budget doing fewer things properly, paced so that it lasts the year it was planned for. The measure of success is not a lower invoice. It is a partner who can look at the invoice and explain it.
There is a reason this practice sits naturally beside our work on judgment. Cost discipline and analytical discipline are the same habit applied to different resources: knowing what a piece of work is for before commissioning it, and being able to defend it afterwards. A token P&L the partners can read is not a technical artefact; it is a governance document, and a firm where judgment is the product should insist on having one. If your firm's AI spend has become a number nobody can explain, we are glad to talk it through.