This month, Amazon quietly switched off a leaderboard. The board was called Kirorank, and it scored employees on how much they used the company’s internal AI tools. It was switched off because employees had started assigning autonomous agents to do work that did not need doing — manufacturing activity for no reason other than to climb the rankings. Dave Treadwell, an Amazon senior vice-president, put it to staff plainly: please don’t use AI just for the sake of using AI.
One of the most sophisticated technology companies on earth had to formally instruct its own engineers to stop using AI pointlessly.
Amazon is not an outlier. Meta ran a near-identical experiment, complete with a leaderboard and status titles for the heaviest users. It registered more than sixty trillion tokens in thirty days before being shut down within forty-eight hours of the data leaking. Microsoft has reported comparable dynamics. Three of the largest technology companies in the world, one failure mode, all surfacing in the same few weeks. The phenomenon now has a name — tokenmaxxing — and it deserves the attention of every executive who has signed off on an AI adoption target this year.
The timing sharpens the point. This week, as the leaderboards came down, Anthropic released Claude Fable 5 — the first publicly available model in its Mythos class, priced at $10 per million input tokens and $50 per million output tokens, and able to work autonomously for longer than any model the company has shipped. From 23 June it leaves Anthropic’s flat-fee subscription tiers and moves to usage credits. Hold those three facts together. The most capable models are now the most expensive per token, the most autonomous per task, and increasingly billed by consumption. Every one of those vectors raises the price of measuring the wrong thing.
Focus On: When measure becomes the target
The instinct is to file tokenmaxxing under “AI gone wrong.” That reading is comfortable and entirely mistaken. What we are watching is a management failure that was diagnosed half a century ago — twice, in the same year.
In 1975, the economist Charles Goodhart gave us the principle that now carries his name: when a measure becomes a target, it ceases to be a good measure. The same year, an organisational psychologist named Steven Kerr published a paper in the Academy of Management Journal with a title that reads like a verdict on Amazon’s leaderboard: “On the Folly of Rewarding A, While Hoping for B”. Kerr’s argument was that organisations habitually build reward systems that pay off handsomely for one behaviour while the people running them hope, often sincerely, for a different one — and are then bewildered to get the behaviour they paid for rather than the one they wanted. Goodhart explains why the metric rots. Kerr explains why management keeps building the rotten metric and acts surprised at the result.
Every executive has watched both operate. Sales teams that hit call-volume targets by rushing customers off the phone. Support agents who close tickets fast and resolve nothing. Universities, in Kerr’s own example, that reward research and publication while hoping for good teaching, and get precisely what they pay for. The mechanism never changes. People optimise for what is counted and rewarded, not for what is quietly hoped for, and the metric detaches from the outcome it was meant to stand in for.
The measurement gap is the accountability gap
The numbers around enterprise AI are, on inspection, unsettling. A July 2025 study from MIT’s Project NANDA found that ninety-five per cent of enterprise generative AI pilots failed to deliver measurable return. Forrester reports that only fifteen per cent of AI decision-makers saw an EBITDA lift over the past twelve months, and fewer than a third can connect AI to any change in the profit-and-loss account at all.
The more revealing data point comes from Grant Thornton’s 2026 AI Impact Survey of 950 senior leaders. Organisations with fully integrated AI were close to four times more likely to report AI-driven revenue growth — fifty-eight per cent against fifteen — than those still piloting. The advantage did not belong to the companies with the most usage. It belonged to the ones who had connected AI to a real outcome. And yet seventy-eight per cent of those same executives admitted they could not pass an independent AI governance audit.
That is the gap that matters. Not the gap between leaders and laggards, but the gap between activity and accountability. Analytics from Jellyfish, covering twelve thousand developers across two hundred companies, makes the point concrete: heavier token consumption did correlate with more output, but the cost per useful unit of work climbed sharply at the top end. More tokens bought less and less. A usage dashboard cannot see any of this. It records the pulse and calls it a diagnosis.
Why rational people built an irrational system
It would be easy to treat tokenmaxxing as a failure of individual judgment. It is the opposite. It is what you get when people respond rationally to the incentives placed in front of them.
Follow the chain. The four hyperscalers — Amazon, Microsoft, Alphabet, and Meta — are tracking combined capital expenditure of between $650bn and $700bn this year, with Wall Street projections exceeding a trillion dollars for 2027, up from under $400bn in 2025. Amazon alone expects to spend around $200bn, the bulk of it on AI and data-centre infrastructure. That scale of commitment rests on a single assumption: that enterprise demand for AI is real, large, and growing. Boards want proof. Executives convert board pressure into departmental targets — at Amazon, more than eighty per cent of developers mandated to use AI weekly. Managers convert targets into usage quotas. Employees, watching managers watch the dashboard, do precisely what the dashboard rewards. No one in that chain is behaving foolishly. The system is failing while every actor inside it is being sensible. That is what makes it dangerous.
This is Kerr’s point arriving fifty years early. The engineer who sets an agent loose on a pointless task to climb Kirorank is neither lazy nor dishonest; the reward system has made that the rational thing to do. Which is what makes Treadwell’s plea — don’t use AI just for the sake of using AI — so revealing. It is an instruction to stop doing A and start doing B, issued to people who are being measured, ranked, and watched on A. Kerr could have drafted the memo for him. He also, in 1975, explained exactly why it would not work.
The antidote already exists
Amazon, to its credit, has begun the correction. It is shifting away from raw token consumption toward what it calls normalised deployments — evidence that engineers are regularly using AI to produce genuinely useful code. That is the right direction of travel, and it points to the broader fix.
The discipline is to measure outcomes that cannot be inflated by activity. I have argued in this newsletter before for a shared portfolio of golden metrics — no more than five measures, each tied directly to the company’s growth algorithm, owned jointly with the CFO. Merged code that ships. Cases resolved without escalation. Hours returned to revenue-generating work. Process metrics such as usage still have a diagnostic role; the error is letting them masquerade as outcome metrics on a board slide. The governance structure has to publish the outcome alongside the input, so the two cannot drift apart unnoticed, and it has to put every adoption leaderboard on a sunset clock. Goodhart’s Law does not care that your dashboard worked last quarter. Decay begins on the next reporting cycle.
The stake nobody is pricing
Here is the uncomfortable implication. The internal adoption data that executives use to justify continued AI investment is the same data that is being gamed. If the demand signal feeding a trillion-dollar buildout is partly manufactured, then capital is being allocated against noise. Boards are approving infrastructure on the strength of adoption curves that may be, in part, fiction.
This will not stay an internal HR curiosity. The economics are tightening from both ends at once. The day Anthropic released Fable 5, it also confirmed the model would move off flat-fee subscription tiers and onto usage credits within a fortnight — the industry-wide drift from flat fees to consumption pricing that has already raised costs for heavy enterprise users, Amazon among them. Under flat fees, tokenmaxxing was waste. Under consumption pricing, it is a direct, metered invoice for nothing. As the practice moves into mainstream coverage, analysts and investors will stop asking how fast adoption is rising and start asking how real it is. The companies that have already replaced activity theatre with outcome accountability will answer that question with evidence. The rest will discover, on a P&L and in front of an audit committee, that usage and value were never the same thing.
So, three decisions worth making soon:
Audit the metric. If your primary measure of AI adoption is usage volume or token consumption, assume it is already contaminated and commission a governance review before it reaches a board slide.
Replace activity targets with outcome contracts. Define what good looks like in units of business output, not units of tool engagement, and pair every input metric with the outcome it is supposed to produce.
Get ahead of the demand-quality question. The infrastructure narrative depends on credible adoption signals. When the market starts probing signal quality rather than signal volume, you want to be the organisation that already knew the difference.
Is your organisation measuring AI adoption, or AI value? And if you switched off the activity dashboard tomorrow, would you still know whether any of this is working?
Follow me
That’s all for this week. To keep up with the latest in generative AI and its relevance to your digital transformation programs, follow me on LinkedIn or subscribe to this newsletter.
Disclaimer: The views and opinions expressed in Chronicles of Change and on my social media accounts are my own and do not necessarily reflect the official policy or position of S&P Global.
