Skip to main content

Benjamin
Charity

Published: August 16, 2026

Usage Is Not Value: Measuring What AI Actually Changed

Reading time: 7min

Every AI rollout eventually produces a dashboard, and the number on it is almost always some form of usage. Tokens consumed, active seats, messages sent, percentage of engineers with the assistant enabled, count of AI-authored pull requests. The number goes up and to the right, which feels like progress, and it gets reported upward as if it were the return on the investment. It is not. Usage is the easiest thing to count and the least informative thing to know.

A retro split-flap counter labeled TOKENS on a black office wall displays a number in the tens of millions while coworkers walk past in a motion blur

Usage answers the wrong question

Usage tells you that a tool was used. It doesn't tell you the tool helped. Those are different claims, and most of the money gets spent without anyone checking the difference. A high token count is consistent with a team getting dramatically more done. It is equally consistent with a team generating more output that then needs more review, more rework, and more cleanup than it saved. The counter can't tell the two apart, because it's measuring activity, and activity was never the goal.

This isn't a new mistake. It's the same one we made with lines of code and commit counts. The measure is tempting because it always moves: give people a tool and usage goes up. A number that moves no matter what happened isn't a signal.

Measure the outcome, not the activity

The honest question is whether the work that matters changed. That pushes you back onto the outcome measures you already trust, the ones that were true before AI and stayed true. On the delivery side, that includes whether you are shipping faster without shipping more defects. Change failure rate is the one to watch, because the failure mode of AI-accelerated work is producing more, faster, with more mistakes riding along. If throughput rose and change failure rate rose with it, you didn't get faster. You found a faster way to create rework.

Then ask the question usage can never answer: what work stopped happening, and was it the right work to stop. The point of the tool is not to add activity on top of everything people already did. It is to remove or compress work that used to cost time, and free that capacity for something higher value. If nobody can name what stopped, or the freed time went straight into producing more of the same low-value output, the tool changed the inputs and not the result.

The waste the counter rewards

There's a second thing the counter can't see, and it shows up even when every team is moving a real goal. Put twenty people on twenty projects and look at the first third of each one. It's largely the same work: the same discovery, the same evaluation of approaches, the same prompts, scaffolding, and decisions about how to start. Each person's usage is defensible on its own, because their project genuinely needed that work. Across the org, though, nineteen of them are re-deriving something a colleague already figured out, and the token counter records every duplicated pass as productive activity. It's worse than a blind spot. Duplication makes the usage chart look better.

This is why capture matters as much as measurement. The decisions, templates, prompts, and half-built artifacts that come out of that first stretch are worth more than the output of any single project, because they move the starting line for everyone who comes after. An org that captures and shares them starts each new project a third of the way in. An org that doesn't pays for the same first mile twenty times and reports the bill as adoption. So add one more question to the review: what did this project learn that the next one won't have to relearn, and where does that live.

The trap on the other side

There are two ways to get measurement wrong, and vanity metrics are only one of them. The other is to measure nothing, to wave off the question because knowledge work is hard to quantify and run the whole program on vibes and anecdote. That fails differently and just as badly, because you can't tell a real gain from an expensive habit, and you keep funding both.

The over-correction from there is to build an elaborate ROI apparatus, instrumenting every task and attaching a dollar figure to every saved minute, which costs more to run than it ever returns and produces numbers that look precise but that nobody actually believes. The useful position sits between them. Track a small number of outcome signals you already believe, connect the usage to whether those moved, and resist the urge to manufacture precision you don't have.

A short way to tell if AI created value

Three questions get you most of the way, without a new measurement program:

  • Did a delivery or business outcome you already track actually move, and did it move without a matching rise in failures or rework.
  • Can you name the work that stopped happening, and did the freed capacity go to something higher value rather than just more output.
  • If you turned the tool off tomorrow, would you feel the loss in a result, or only in the usage chart.

If the only honest answer is that the usage chart would drop, you've been measuring a habit, not value.

Where usage metrics are fine

None of this makes usage useless. During an initial rollout, reach is a legitimate thing to track, because before a tool can create value it has to actually be adopted, and usage tells you whether that is happening. The error is not looking at usage. It is calling it value, and letting the adoption number stand in for an impact the tool has not yet been shown to produce. Track usage as adoption, label it as adoption, and don't let it graduate into an ROI claim it can't support.

Where people will push back

"Usage is a leading indicator of value."

It can be, but only if you eventually connect it to a lagging outcome. A leading indicator you never tie to a result isn't leading anywhere. It's a number that reliably goes up. The discipline is to treat usage as a hypothesis about value and then check whether the value showed up. If you never run the second half, you don't have a leading indicator. You have a number you enjoy reporting.

"You cannot measure knowledge-work productivity."

Not perfectly, and you don't need to. The claim that it is unmeasurable usually smuggles in the conclusion that you should stop trying, which leaves you measuring activity by default, which is worse than an imperfect outcome measure. You can detect whether the outcomes you care about moved. That is enough to steer by, and it is more than a token count will ever give you.

"Leadership wants the adoption number."

Then give it to them, clearly labeled as adoption. There's nothing wrong with the number itself. The trouble starts when the organization celebrates the token count as if it were impact, and makes decisions as though usage and value were the same thing. Report the reach and the outcome side by side, so the activity is never mistaken for the result.

The takeaway

AI usage is easy to measure, which is exactly why it fills dashboards and why it misleads. It proves the tool was used, not that it helped, and those aren't the same claim. Usage is not value. It is the cost you are hoping turns into value. To know if it did, look at whether the outcomes you already trust actually moved, and at what work stopped and whether it should have. Everything else is a number that goes up on its own.

Further reading

Build, Scale, Succeed

Join others receiving expert advice on
engineering and product development.

Newsletter Subscription

No data sharing. Unsubscribe at any time.