How we work / Spend / Tokenomics
Tokenomics
Token economics is a FinOps method for measuring what AI costs and what it is worth, one token at a time. Turn the rest of your technology spend into standard units the same way, and tracking it day by day, week by week or quarter by quarter becomes simple.
Token economics is the FinOps Foundation's term and method, not ours. We apply it, alongside the FinOps Framework, FOCUS, ISO/IEC 19770 and TBM.
What a token is, and what token economics means
A token is the small unit of text, image or sound an AI model reads and writes: often a piece of a word. The FinOps Foundation gives a sense of scale: about 1,500 English words come to roughly 2,048 tokens, though every model counts a little differently (FinOps Foundation, "Token Economics: The Atomic Unit of AI Value", 10/05/2026, read 04/10/2026). AI providers bill by the token, and they price the tokens you send in and the tokens the model sends back separately, often at different rates.
Token economics, or tokenomics, is what the FinOps Foundation calls FinOps applied to AI
: metering AI use by the token, tying it to the team, product or customer that caused it, and connecting it to the value it creates (same source).
Fine-grained control of AI spend
A monthly AI bill tells you the total. Tokens tell you where it came from. Once every request is metered, the same bill can be read many ways:
- Cost per request: what one question, one summary or one agent step costs.
- Cost per feature: which part of your product or workflow uses the most.
- Cost per customer: what it costs to serve each customer, and whether your price covers it.
- Cost per team: who is using what, shown back to them in the same units.
- Cost per outcome: the measure that matters most, such as cost per case resolved or per document processed.
The FinOps Foundation describes this as a path: early AI unit economics often starts with cost per token, then grows toward outcome measures such as cost per assist, per agent action or per case deflected (FinOps Framework, Unit Economics capability, read 04/10/2026).
The detail also shows where the money goes: long system instructions sent with every request, more context than the task needs, a larger model than the job requires, long answers, and retries. Each is something your teams can change.
A worked example
An illustrative example. All figures, including the prices, are made up to show the idea; they are not any provider's prices.
A support assistant handles 200,000 requests a month. Each request sends about 1,500 tokens in (the question, its instructions and some history) and gets about 300 tokens back. Say input tokens cost USD 2 per million and output tokens USD 8 per million.
| Tokens a month | Cost a month | |
|---|---|---|
| In | 300 million | USD 600 |
| Out | 60 million | USD 480 |
| Total | 360 million | USD 1,080 |
That is about USD 0.0054 per request. If the assistant resolves 40,000 cases a month, it costs about USD 0.027 per case resolved, which is the number to set against what a case costs any other way.
Now the team trims 500 tokens of instructions that every request carries. Input drops to 200 million tokens, the bill to USD 880 a month, and the cost per case to USD 0.022. Nobody had to guess: the change and its effect show up in the same units the next day.
The same idea across the whole estate
Tokens work because they turn a messy bill into a standard unit with a rate. The rest of your technology spend can be treated the same way: every part of the estate gets its own standard spend unit, with an agreed rate that includes what it really costs to run.
| Part of the estate | A standard spend unit | What the rate includes |
|---|---|---|
| AI | per 1,000 tokens, or per request | provider charges, plus the platform around it |
| Cloud | per server-hour, per GB stored a month | the provider's charges after discounts |
| SaaS | per active seat a month | the licence, plus admin time |
| On-premise | per server a month | write-down, power, space, support and people |
| People | per hour of engineering or support time | salary and overheads |
The FinOps Foundation's own list of unit metrics already mixes these: cost per GB stored, per vCPU, per token and per seat alongside cost per customer and cost to serve (Unit Economics capability, read 04/10/2026). FOCUS, its standard layout for bills, now covers AI, cloud, SaaS and data centre costs, so the units can be counted the same way everywhere (FOCUS, read 04/10/2026).
Why tracking becomes simple
Once usage is counted in standard units and each unit has a rate, spend is just units used times the rate. The same table can be read at any time scale:
- Daily: spot a jump the day it happens, and see straight away whether more units were used or a rate changed.
- Weekly: each team sees its own units and cost, in terms it recognises.
- Monthly: show back costs fairly, by use rather than by headcount or habit.
- Quarterly: forecast from expected units, set budgets, and check prices against the real cost to serve.
Without standard units, every one of those questions means rebuilding the numbers from bills that each arrive in a different shape.
The honest part: the effort up front
This takes real work at the start, and it is worth saying so plainly:
- Every cost has to be found and given an owner, including AI on personal cards and AI built into software you already pay for.
- Finance and technology have to agree the units and the rates, including people's time and the full running cost of equipment you own.
- Usage data has to be switched on and connected, and AI requests labelled by team, feature or customer.
- Some token costs are hidden. Where AI is built into a product you pay for by the seat, the vendor manages the tokens and you cannot see them; the FinOps Foundation advises measuring the value each seat delivers instead (FinOps Foundation, Tokenomics in SaaS, read 04/10/2026).
For one focused area this is a matter of weeks, not days; our own setups for smaller organisations run from about 4 weeks (cloud or AI) to about 7 (the whole technology estate).
Why it pays back: once it is in place, questions that used to take a spreadsheet and a week take minutes. Jumps are caught the day they happen. Forecasts are built from units, not last year's total. Prices can be set against the real cost to serve. And every team is judged on its own use.
Start small
The FinOps Foundation's advice applies here as everywhere: start small and grow as the value shows, through its Crawl, Walk, Run stages (FinOps Foundation, read 04/10/2026).
- 1
Crawl: one AI use, one team
Meter one AI use case by the token and work out its cost per request. Agree one outcome to measure it against.
- 2
Walk: more uses, more units
Add AI uses and the first non-AI units (seats, storage, server-hours). Show each team its costs monthly.
- 3
Run: the whole estate
Standard units and rates across AI, cloud, SaaS, on-premise and people, tracked from daily to quarterly, feeding budgets and prices.
Governance underneath
Token economics only works on top of governance. A token can be tied to a team only if every AI use has an owner and its own account or key; tokens spent through shadow AI, on tools nobody approved, cannot be traced to anyone. The register of systems, owners and access that governance builds is the same register tokenomics counts against.
Read why governance comes first, and how spend observability and optimisation build on it.
Check how ready you are
Our free spend self-check includes a set of questions on token readiness: whether you can see AI use by the token, tie it to teams and features, and track costs per unit over time. About 20 minutes, no sign-up.
How we help
AI spend management builds the cost model per request and per use case, with tracking and alerts your teams switch on. Technology spend management sets the units and rates across the whole estate. Typical prices for both are on our Pricing page.
Start with a conversation.
One call to understand what you spend on and what worries you. If an engagement fits, you get a written scope and price.
Book a first call