Articles

    Token costs are a design decision, not a skills problem

    The question reaches every leadership team eventually. First AI was an experiment, then it was a licence, and at some point it becomes a line in the accounts growing faster than anyone budgeted for. That is when someone asks what this is costing us. The usual answers are skills answers. Set caps and ownership. Hire that rare person who understands both cloud cost and AI architecture. Be a better engineer and know where not to use a language model. All three are right. All three also move cost control inside somebody's head, which is why none of them scales.

    Four layers, one name

    The first reason these conversations stall is not technical but conceptual. The word covers four different things. Individual subscriptions: the cost arises from a monthly fee per user, it is raised by experts and line managers, and it is driven by personal habit. Organisational API usage: the cost arises from volume run and from model and context choices, it is raised by IT, architects and finance, and it is driven by implementation choices. Own compute: the cost arises from capital outlay and utilisation, it is raised by infrastructure and security, and it is driven by the stability of volume. Total impact: the above plus integration, data work and human time inside the process. It is raised by executives, and it is driven by process structure. These are independent axes. You can run open weights on a hyperscaler, a closed model over an API from a European data centre, or your own weights on your own hardware. Each combination carries different economics and different risk. Collapse them into one question and every answer is wrong at one of the layers. One practical note: owning compute wins only at high and steady utilisation. Agentic processes fire from events and run in bursts, so usage-based cost wins until volume is both large and stable. Below that threshold, the lever sits elsewhere.

    A number without a denominator is not a number

    A monthly spend equal to several expert salaries is cheap if it replaces the work of several experts. The same spend is expensive if nobody can say where it went. The question is not how high the cost is, but whether it has a denominator. In an individual-led model, there is none, and there cannot be. Cost grows with headcount and habit and has no structural ceiling, because a person can always ask again. The cost exists, and the output exists, but nothing readable connects them. In an agentic process, the denominator is the run. The process fires from an event, follows a defined path and produces a known result. Attributable cost can be managed. The harder question follows immediately: what do you divide the run by? If the process replaces repetitive manual work, the denominator is hours or headcount. If it produces bids, it may be a won bid. If it reduces defects, it is a complaint or a rework cycle. If it shortens throughput, it is calendar time, which often matters most to executives and is measured least. That choice is not technical. It states what you actually expect from AI, and it differs from process to process.

    The cheapest run is not cheap if the output is wrong

    Here is the part that almost always goes unnoticed. A cheap run producing a mediocre bid costs more than an expensive run producing the right one. The difference appears in no metric at all, because an experienced person quietly fixes the output on its way through. That correction work is the real cost of the process, and it is invisible in a cost report. It is not a line on an invoice. It is hours somebody spent turning an inadequate output into an adequate one. The cause is not the price of the model but missing definition. If the acceptance criterion has not been written down, the machine cannot hit it, and people cannot judge it consistently. In workshops, this is the part that looks quick and turns out to be hard: once a group starts discussing what a good result actually is, it often emerges that no shared view exists, and that everyone has been correcting the work in their own direction for years. Acceptance criteria are therefore not a footnote to the cost question. They are what makes cost calculable.

    What if you cannot calculate the value

    There is a fair objection here. Most companies carry significant costs whose return is unknown or impossible to establish reliably, and cost accounting rests on assumptions and rules of thumb anyway. Why should AI be the exception? It should not be. But prioritisation does not need an exact figure. It needs an order. In large definition programmes, we use an estimate built from three multipliers, and the group doing the work makes it, not an outsider. How often the work is done, for example, project starts per year. How long one instance takes. How many people do the same work the same way. Together, these give an order-of-magnitude estimate, iterated once. It is not precise, but it is honest, and that is a different thing from a supplier's benefit projection. The decisive finding is the shape of the distribution. When every workflow in an expert organisation is estimated the same way, impact does not spread evenly. In our observations, roughly a third of the workflows account for around 90 per cent of the estimated total impact. That has a consequence which answers the objection in full. Halve every estimate and the top of the list does not change. Divide them all by five, and it still does not change. The error is large relative to the number and small relative to the ranking, and ranking is what the first decision needs. The reason lies in where workflows land. They do not spread evenly along the scale; they cluster into groups separated by an order of magnitude, with roughly two orders of magnitude between the smallest and the largest. A twofold estimation error moves nothing between groups, because the gaps are far wider than the error. This is worth saying plainly to a buyer, because it is an unusual promise. We do not promise the estimates are right. We promise they are in the right order of magnitude, and an order of magnitude is enough to decide where the first investment goes. The reason for the skew is worth stating, because it is a business finding rather than an AI one. An expert firm typically has a long tail of bespoke small assignments carried out by one or two people, and at the other end a few large pieces of work done the same way by dozens. The tail is not worthless. What it lacks is productisation and volume, and that is a strategy question that existed before AI. Definition is simply the first thing that forced it to be measured.

    The answer is in the definition, not in the expertise

    Back to the start. The third answer was right: the most expensive token is the one burned on a task a few lines of code would have handled. In the majority of steps, there is nothing to reason about. Retrieval, formatting, checking against a rule, moving data between systems. A rule is enough, or a small model is enough. But as long as the choice of model depends on who happened to build the pipeline, it is craft work. Quality varies with the person; it does not carry over to the next process, and across a portfolio of thirty processes it does not scale. When the same choice is written into the specification step by step, it gets reviewed, it can be audited, and it repeats identically across the portfolio. The principle matches the one governing autonomy levels: choose the lowest sufficient level, not the highest available. A good engineer knows where not to use a model. A specification knows it without them. Cost, then, is not primarily a question of expertise. It is a question of definition, and it is settled before the first run, not on the first invoice.

    Next step

    Start with a 30-minute discovery conversation. We identify where your expert work concentrates and assess which process is worth defining first. No implementation commitment.

    Book a conversation

    Frequently asked questions

    Q:Aren't spend caps and reporting enough?

    They prevent surprises but say nothing about what the money bought. A cap limits spend; a denominator makes it manageable. You need both, and only the second produces decisions.

    Q:Should we run models on our own hardware?

    Only at high and steady utilisation. Agentic processes run event-driven and in bursts, so for most expert organisations, usage-based cost stays cheaper for a long time. Settle the question once volume is known.

    Q:How do we find out what one run costs?

    It is not a measurement problem. If you cannot read the cost of a run, the process has not been defined to be run: it starts from a person rather than an event, and no single run can be bounded. Definition produces the metric as a by-product.

    Q:We don't know what a good result looks like. Can we still start?

    Yes, and that is the most common starting point. Writing down the acceptance criterion is part of definition, not a prerequisite for it. It is often the stage from which an organisation gets the most value before a single agent is built.