As many know, my formation is law, not mathematics. The last time I sat an exam involving calculus, Italy still had the lira. So I bought Anil Ananthaswamy’s Why Machines Learn in July as a Summer reading. Turns out this surprisingly approachable book teaches us a lot about how machines learn, allowing us to put things into perspective.
This week, I’d like to share some reflections on it.
Focus On: What the book gives a non-mathematician
The first is a chronology that reorders your sense of the field. Legendre published the method of least squares in 1805, to fix the orbits of comets. Rosenblatt built the perceptron in 1958. Backpropagation reached neural networks in 1986. Hopfield and Hinton collected a Nobel Prize in physics in 2024 for work finished four decades earlier. Almost nothing in the mathematics is recent. Lampedusa has it that everything must change so that everything can stay the same; this ran the other way round. The mathematics stayed still, and data, silicon, and eventually electricity moved around it.
The second is that Ananthaswamy has solved the problem of the equations. They are all there, worked properly, for readers who want them. He also builds every chapter so that a reader who skips them still arrives at the right destination. I skipped a good number. I understood the shape of every idea in the book, and I could not reproduce a derivation from it, and I have made my peace with that. It is written for people who make decisions about these systems without building them, which is most of us.
What survives is four sentences. Learning means fitting a function to examples rather than writing rules, so the programmer supplies the examples and gives up the algorithm. To make that possible the world is converted into geometry: contracts, customers, images, and sounds all become points in a space of many dimensions. Fitting is optimisation — predict, measure the error, adjust, repeat some billions of times. And the object you end up with approximates the distribution it was fitted to, and nothing else.
That last sentence has cost me some sleep. Three arguments I have made in this newsletter turn out to sit underneath it, and in each case I had reached the conclusion by a longer and less reliable route.
Hallucination. I have argued that fabrication is structural rather than a defect, and I got there through the training incentives: grading schemes that reward a confident answer over an admission of ignorance. That reasoning holds. It is also not the deepest layer. A generative model estimates a probability distribution over what comes next and draws from it, which means a hallucination is a perfectly well-formed sample from a distribution that happens not to match the world. The system did exactly what it is. I had the right conclusion approached from one floor up.
Cost. One week I argued that the price of running an agent is a distribution rather than a number, and that management accounting has no instrument for that. I reached it through a KPMG survey, an arXiv study, and Uber’s overspent budget. Ananthaswamy would have handed me the same conclusion in a sentence: a system that samples from a distribution returns a distribution of outputs, and a distribution of outputs carries a distribution of costs. Being right by invoice when you could have been right by first principles is a particular kind of humbling.
The moat. The chapter on support vector machines contains the one piece of mathematics I can now explain over dinner, and it is the idea I would put in front of a chief executive if I were allowed a single page. Some problems cannot be separated cleanly in the coordinates you are given, and become straightforward in a different set. You do not need a cleverer boundary. You need a different space.
I have argued for a while that the deployment layer, not the model, is where advantage accumulates, and I argued it from watching deployments succeed and fail. The kernel trick tells me why. Your taxonomy, your schema, your context pipeline, the metadata nobody wanted to own: those constitute the coordinate system the model reasons in. Menlo Ventures has enterprise open-weight share falling from 19 per cent to 11 per cent, which reads less as open source losing than as buyers concluding the model is not the variable. Nasuni surveyed a thousand enterprise buyers in May and found 94 per cent struggling with unstructured data while 59 per cent named AI their leading investment. Those firms are shopping for better boundaries and declining to change the space, which is the pattern I was arguing for at the start of the month without knowing it had a name in 1992.
The part I hadn’t seen coming
If a model approximates the distribution it was fitted to, its competence has an edge. Obvious, once stated. What I had never registered is that every capability number I quote is measured in the middle. I wrote five days ago that we are spending real money on systems we cannot instrument. The problem is worse than I put it: the instruments exist, and they are pointed at the wrong part of the distribution.
OpenAI’s GDPval is serious work: 1,320 tasks, 44 occupations, output rated as good as or better than industry experts on roughly half, performance tripled in a year. Then read the exclusions. No ambiguous briefs. No tasks requiring clarification. No iteration, no client feedback, no accumulated context. METR is blunter about its own suite, conceding it is “much cleaner than real economically valuable labor”.
Neither organisation is concealing anything. The failure is at my end, and at the board’s. We size multi-year commitments against measurements taken on the well-specified middle, then deploy into the ambiguous, contested tail, which is the only part of the business anyone is paid for. Distinctiveness is thinly represented in any distribution by definition. So capability and value run in opposite directions, and no quantity of capital expenditure reverses a fact about the geometry.
The failure data has already moved. ChatSee’s analysis of ten thousand enterprise AI failure events puts hallucination below ten per cent of the total, escalation and resolution breakdowns at 31 per cent, and execution failures up 62 per cent while hallucinations fell. Treat the precise numbers with the caution any vendor’s research deserves; the direction is not in doubt, and it is the direction the mathematics predicts.
Which has changed the question I ask in meetings. Not how accurate the system is. How often it correctly recognises that it is out of its depth and hands over. Nobody I have asked can state that number, and it belongs on the short list of measures your vendor has quietly stopped reporting.
The next enterprise AI failure serious enough to reach a front page will not be a fabricated citation. It will be an escalation that never happened, on an input the system had never seen, and the post-mortem will find everything performed exactly as the mathematics guaranteed. No model card will have anticipated it, because no model card is measured on your edges. Someone in your organisation should be able to name those edges by the end of the quarter.
Read the book if that person is you.
Follow me
That’s all for this week. To keep up with the latest in generative AI and its relevance to your digital transformation programs, follow me on LinkedIn or subscribe to this newsletter.
Disclaimer: The views and opinions expressed in Chronicles of Change and on my social media accounts are my own and do not necessarily reflect the official policy or position of S&P Global.
