On 4 June, Anthropic published a number built to be quoted. Its engineers, the company said, now ship eight times as much code per quarter as they did across 2021 to 2025. The figure travelled exactly as intended — into headlines, board decks, and the long-running argument over whether AI has finally earned its capital. What travelled less well was the sentence Anthropic set beside it in its own report: lines of code is an imperfect measure, the company conceded, one that counts quantity over quality, so the figure is “almost certainly an overstatement of the true productivity gain.”
A vendor disowning its own headline metric in the same document is unusual. What happened next is more instructive. The executive who runs the team that produced the number spent the following week on a podcast explaining why she doesn’t trust it.
Focus On: The Number Operators Stopped Believing
Fiona Fung leads Claude Code and Cowork at Anthropic. Asked how she now measures her engineers, she went through the throughput proxies and rejected them one by one: lines of code, commit counts, tokens burned, the “token-maxxing” that turns model usage into a leaderboard. Her phrase for it was exact: don’t forsake motion for progress. A team optimising output volume, she argued, has built a faster treadmill and mistaken it for ground covered.
Read that as expertise, not modesty. She is watching the constraint move in real time. When Claude began writing more than 80% of Anthropic’s merged code, the bottleneck didn’t disappear; it moved. Production stopped being scarce; verification became the binding limit. Anthropic names the mechanism in the same report: Amdahl’s law, the principle that accelerating one stage of a process only exposes the next slowest one. Generate code eight times faster and you haven’t finished eight times sooner. You’ve created a review queue no human team can clear by hand, and made trust — not typing — the rate-determining step.
The lesson travels well beyond one lab, because the asymmetry it exposes is now general. A metric informs only when the thing it counts is scarce. Lines of code meant something when a person wrote each one: the count proxied effort, and effort was the constraint. Remove the constraint and the proxy goes hollow. The same holds for every volume measure currently sold to the C-suite as proof of AI’s return: tokens consumed, content variations produced, documents summarised, seats logged in. Each counts activity that has become almost free, which is exactly why each has stopped carrying information.
The trap is that volume still feels like progress, even to the people generating it. The cleanest evidence comes from outside the vendors. In a randomised controlled trial, the research group METR set experienced developers loose on real tasks in their own repositories, half permitted AI tools and half not. They forecast that AI would make them 24% faster. Afterwards, they reported feeling 20% faster. They were in fact 19% slower. A gap of nearly forty points had opened between perceived and actual productivity — and it survived the experience itself. Developers who had just been slowed down came away certain they’d been sped up.
That is the precise failure mode of governing AI by output volume. The number climbs, the dashboard greens, the felt sense of acceleration confirms both, and none of it touches whether the work improved or the outcome moved. As I argued in The Great Disconnect, this is Solow’s paradox reproduced at company scale: AI visible everywhere except in the results that matter. The macroeconomic version reads as flat total factor productivity. The organisational version reads as a quarter of record tool usage and no discernible change in what the function delivered.
The Decision Underneath the Dashboard
The question for anyone signing AI invoices, then, isn’t how much output the tools have unlocked. It’s whether you’re measuring the thing your vendor’s marketing department measures, or the thing your vendor’s engineering leaders measure. They’re not the same number, and the distance between them is the whole point.
Another dashboard won’t correct it. What counts as the unit of value has to change. Stop measuring output and start measuring verified output: work that survives review and delivers what it was commissioned to do. The ratio that matters is the share of generated volume that clears the bar without a human having to redo it. A team producing ten times the drafts while that share falls hasn’t grown more productive; it has shifted labour from creation to correction and renamed the move efficiency.
This sharpens a discipline I’ve set out before. In The AI Dividend Trap I argued for a metrics contract with finance that can’t be easily gamed. The newest way to game it is now in plain view: adopt your vendor’s headline throughput figure as your own KPI. Goodhart’s law does the rest: a measure that becomes a target stops being a good measure. The figure arrives pre-validated and points in the flattering direction, and it counts precisely the quantity that has stopped being scarce. A finance-grade operating model has to exclude it by design, treating raw volume as an input to be minimised per unit of verified outcome, never as an outcome itself.
There’s a tell that separates the two postures. If your AI dashboard went dark tomorrow — every usage chart, every token count, every adoption percentage — would one business outcome change? If the honest answer is no, you’ve been measuring the treadmill and calling it the journey.
Anthropic, to its credit, said the quiet part in its own footnote. The firms building these systems have already moved their internal attention from how much their models produce to how much of it can be trusted. The organisations buying those systems are, in the main, still applauding the volume. Whoever closes that gap first — whoever learns to measure verification rather than production — converts an impressive chart into a real advantage. Everyone else is left with the chart.
Follow me
That’s all for this week. To keep up with the latest in generative AI and its relevance to your digital transformation programs, follow me on LinkedIn or subscribe to this newsletter.
Disclaimer: The views and opinions expressed in Chronicles of Change and on my social media accounts are my own and do not necessarily reflect the official policy or position of S&P Global.
