Everyone is benchmarking which model is smartest.
From a finance seat, the more useful question is: what does one unit of work cost — and does that cost hold as you scale?
I’ve spent months running my own AI models and applications, pushing hundreds of millions of tokens through them. The real story turned out to be the economics.
Per million tokens (output), top-tier frontier models cost up to ~$50.
The cheapest open-weight options: $0.60.
That’s up to 80x.
No negotiation or optimisation closes that gap. It’s a different cost structure — and it changes which projects are viable.
On capability: open-weight models hold their own on the majority of tasks a finance function would actually automate
– document extraction
– classification
– reconciliation support
– first-pass analysis
– code assistance
The gap only opens at the edge of the distribution: the hardest few % of problems.
So I treat model choice as procurement, not preference:
Standard tier — open-weight models by default, wherever the task is well understood.
Premium tier — frontier models, reserved for high-complexity, edge-of-distribution work.
That routing discipline is how I’ve scaled my AI workload 5–7x over the last few months — while expanding the scope of applications running on it. Hundreds of millions of tokens. Dollar-spend still well under control.
The finance takeaway: inference is a variable cost. And like any variable cost, it responds to design.
The winners won’t be the teams with the biggest AI budgets. They’ll be the ones getting far more work out of the budgets they already have.
I pulled list prices for the most widely used models — input vs output side by side, per 1M tokens, with the cost multiple against the cheapest open-weight option. Sources in the comments.
How is your organisation treating inference: a line item to manage, or a design decision to own?

