Back

Stop counting tokens, start counting outcomes

October 9, 2026

Ask a board member how much the company spends on AI and someone can usually answer by the end of the day. Token consumption is logged, invoiced and easy to put on a chart. Ask what the business received in return and the conversation changes, because those numbers sit in different systems and nobody agreed upfront which ones matter.

Attention is starting to move in that direction. Google Trends data cited by Google Cloud shows four times more interest in “token efficiency” than in “tokenmaxxing” (treating heavy token consumption as a sign of progress) or “token leaderboards” since January 2026. Search interest is a soft signal and we would not build a strategy on it alone. It does suggest that the question is shifting from how much AI an organisation uses to what that use returns.

This article looks at what the latest Google Cloud research says about organisations with accelerating returns, why token spend works better as a cost line than as a headline, and how to build a one page scorecard that a board can read in five minutes. It continues from our article on why most AI initiatives need an orchestrator, not just more models [link to be added], where we argued that a single view of cost and performance is one of the main things coordination provides.

What the Google Cloud research found

Google Cloud and the National Research Group surveyed 2,403 executives across industries and regions. Of those, 84% said they see increasing financial returns from AI initiatives, and 26% of that group reported returns that are accelerating year over year. Google Cloud calls this smaller group “AI ROI Leaders”.

The report identifies three practices the leaders share. Some 48% have set extremely clear ownership and decision making authority for AI and agentic initiatives. Almost half say AI is embedded in core business processes and revenue streams, or is enabling new business models. And 38% run ongoing capability development with required training built into roles.

Agents stand out in the data. In total, 94% of respondents say AI agents contribute to cost savings and revenue, with a slightly heavier weighting towards savings. The impacts cited most often as measurable were faster strategic decision making and increased workforce capacity.

We read this with some care. Google Cloud sells the tools in question, the figures are self reported by executives, and the customer examples come from Google’s own customers. It is a useful signal of direction, not proof. What interests us more than any single percentage is what the three practices have in common. None of them mentions a model or a token count. They describe accountability, placement in the workflow and people’s skills.

Why token spend works as a cost line, not a headline

Token spend measures what goes into the work. It helps with cost control, and we come back to that below. It tells a board nothing about whether the work itself improved, in the same way that a delivery company reporting litres of fuel has told you something real about its costs and nothing about how many parcels arrived on time.

The number can also mislead in both directions. A rising token count might mean an agent is handling far more cases than before. It might equally mean an agent keeps retrying a step that fails, and from the invoice alone you cannot tell which.

In our orchestrator article, we described how organisations often optimise for token consumption because it is the only number they can see easily. That article also cites Gartner’s expectation that more than 40% of agentic AI projects will be cancelled by the end of 2027, with escalating costs, unclear return on investment and weak risk controls as the main causes. A scorecard brings those three causes into view early, while there is still time to adjust.

A one page scorecard for the board

The scorecard works per workflow, not per tool. A board does not need to know which model sits behind a process, but it does need to know whether the process improved. Five lines are usually enough.

One page AI scorecard, built per workflow
Line Question the board is asking Example measure
Ownership Who is accountable for this initiative? A named owner per workflow, with authority to change the process
Workflow outcome Did the work get faster, cheaper or more accurate? Time to first assignment, cycle time for invoice exceptions, resolution rate
Cost per outcome What does one finished unit of work cost now? Total cost of AI and people divided by completed cases, with token spend as one input
Quality and control Can we trust the output, and can we show why? Share of outputs corrected by a person, error rate on a sample, audit trail coverage
Adoption and capability Are the right people using it well? Share of intended users active weekly, completion of required training

Ownership. Each AI initiative needs one owner with authority over the process, not only over the tool. When an agent touches the work of three teams, the owner is the person who can change all three. This lines up with the Google Cloud finding that 48% of the leaders have extremely clear ownership.

Workflow outcome. Choose the one number the workflow exists to improve. In the Dataiku research we referred to in the orchestrator article, invoice exception handling went from three to five days down to minutes once the workflow was orchestrated, and the logistics company Geodis reported 60% faster ticket assignment. Both are measures of the work itself, and neither depends on how many tokens were used.

Cost per outcome. Divide the full cost of the workflow, AI and people together, by the number of completed cases. Token spend then becomes one input to that figure. If token spend doubles while cost per case falls, the board has a good story to tell. If token spend falls while cycle time stays flat, the saving has not yet produced anything.

Quality and control. Track how often a person corrects or overrides the output, the error rate on a sample of cases, and how much of the activity is covered by an audit trail. These numbers matter commercially as much as for compliance. A fast process whose output needs correcting downstream has simply moved the cost to another team.

Adoption and capability. Count how many of the intended users work with the tool each week and whether the required training has been completed. Google Cloud found that 38% of its leaders build training into roles. A workflow that nobody uses produces no outcome, however good the technology behind it.

One caution on the impact the survey cites most often. Increased workforce capacity is real, but it only becomes value once someone decides what the freed hours are for: more cases handled, a hire avoided, a delivery date brought forward. Without that decision it stays a soft number that is hard to defend in front of a CFO. Write the intended use of the freed time into the scorecard from the start.

Some organisations do state value in money. Highmark Health, for example, reports $27.9 million in value in 2025 from its Sidekick assistant, and Elanco reports an estimated $1.9 million ROI. Both figures come from Google Cloud’s customer stories, so we treat them as claims to examine rather than benchmarks to copy.

An illustrative example

To make the mechanics concrete, take a fictional logistics software company with 150 employees and a support team of 20. It introduces an agent that classifies incoming tickets and routes them to the right specialist. The figures below are invented to show how the scorecard reads. They are not results from a real client.

Illustrative figures, invented to show how the scorecard reads. Not results from a real client.
Line Before After 90 days
Owner Shared between support and IT Head of customer support
Median time to first assignment 6 hours 1.5 hours
Cost per resolved ticket €9.20 €7.10, including AI costs
Tickets corrected by a person Not tracked 14%
Support staff using the agent weekly Not applicable 18 of 20

A board member can read this in under a minute. Token spend sits inside the cost per ticket line. In this example it rose, because the agent now touches every ticket, and cost per ticket still fell. The 14% correction rate tells the owner where to improve the routing rules next. None of that is visible from a token invoice.

Setting it up in 30 days

You do not need a measurement programme to start. You need four decisions.

  1. Choose two or three workflows. Pick processes that are already running or about to launch, rather than trying to cover every AI tool in the organisation.
  2. Capture the baseline before launch. Without a starting figure there is no before and after. If the AI is already live, reconstruct the baseline from last quarter’s records in the ticketing, finance or CRM system.
  3. Name an owner for each workflow. Give that person authority to change the process, and make them the one who presents the numbers.
  4. Agree the review rhythm and the decision rules in advance. Monthly for owners and quarterly for the board works well, with a written view of which result would lead to scaling up, adjusting or stopping. This also addresses the cancellation risk Gartner describes, because stopping becomes a planned decision instead of a reaction to a surprise.

Where tokens still belong

Token efficiency is a sensible discipline, and the shift in interest towards it is healthy. Routing simple tasks to a lighter model and keeping the most capable model for complex reasoning keeps cost in proportion to value, as we described in the orchestrator article. In the scorecard, token spend belongs inside the cost per outcome line, where it explains why the figure moved. Used that way it supports the outcome story without replacing it.

Limits worth stating

The scorecard has limits. Survey results like Google Cloud’s reflect what executives report, and the figures quoted from customers are claims the companies themselves make. Ninety days gives a first reading of a workflow, and some value, such as better decisions, takes longer to show and needs proxies that you label honestly as proxies. A scorecard also has its own cost, so keep it to one page per workflow and drop any line that nobody uses to make a decision.

The strategic takeaway

Boards are asking a fair question: what are we getting for this spend? Token counts answer how much. The scorecard answers what. The organisations Google Cloud describes as accelerating share clear ownership, AI placed inside core processes and people trained to use it, which are decisions about how the business works rather than about which model to buy.

We have been building AI systems since 2015, and in our experience the projects that hold up are the ones where somebody can say, on one page, what changed and who is responsible for it. If you are preparing the next board update on AI and want to decide which numbers belong on the page, let’s talk about how this framework applies to your product.

Author

Chairunnisa Irianto

Nisa is a Marketing Manager at Itsavirus, a strategic software development partner working with companies across Europe and Southeast Asia. She writes about AI, application modernisation, and how businesses turn technology into practical results.

Latest insights

A sharp lens on what we’re building and our take on what comes next.

See more
How to build a knowledge base that gets smarter over time with Obsidian and Claude Code
Your AI keeps forgetting. Here's how to stop repeating yourself
OpenClaw is exciting. But, here's what you need to secure before you experiment

Latest insights

A sharp lens on what we’re building and our take on what comes next.

See more
Before AI-generated code reaches customers: a release review checklist
Shadow AI: the tools your team already uses, and who is accountable for them
Switzerland is building a sovereign cloud. Here is why every European business should be paying attention

Latest insights

A sharp lens on what we’re building and our take on what comes next.

See more
The cost of leaving your cloud provider ends on 12 January 2027
Choosing an AI model just stopped being simple, and 25 US tech companies are fighting to keep it that way
Claude Fable 5: Launched, praised, then pulled within 3 days

Latest insights

A sharp lens on what we’re building and our take on what comes next.

See more
Workshop : From Idea to MVP
Webinar: What’s next for NFT’s?
Webinar: finding opportunities in chaos

Latest insights

A sharp lens on what we’re building and our take on what comes next.

See more
How we helped Ecologies to turn survey results into reliable, faster reports using AI
How to deal with 1,000 shiny new tools
Develop AI Integrations with Itsavirus
No items found.