ada könig
← all pieces
Urtext · 2026.08.07

The Model Was Never the Scarce Part

Microsoft now runs token budgets on its own engineers while Europe's integrators book the margin: the scarce part of AI work is the judgment about where the machine goes.

Microsoft has put a meter on its engineers' thinking. Since July its divisions have run against AI token budget targets, individual engineers can watch their own consumption, and executive vice president Jay Parikh told staff that "tokenmaxxing is not what we are optimizing for". Internal guidance concedes that many engineers spend in the range of hundreds of dollars a month, some a few thousand. The company also moved its internal default to GPT-5.6 Sol, the cheaper option.

That memo is the most honest accounting anyone has published on what a token is worth. Microsoft sells the world a Copilot for every developer. Inside its own building it is now deciding which problems deserve the expensive model and which get the cheap one.

The default is the tell. A budget line still leaves room to argue your case; a default is a decision taken once, for everyone, by people who will never meet the problem you are working on. Moving every engineer to a cheaper model sets a ceiling on the work, and it arrives dressed as an efficiency measure.

Agents consume the way they do because one instruction fans out. A single request triggers tool calls, searches, retries, background passes, and a second look at the model's own output, so the counter is measuring attempts. It cannot separate a productive attempt from a flailing one, which means a token budget puts a price on persistence and leaves the quality of the question untouched. Anyone who has attached a profiler to a running system knows what follows: you optimise what the profiler measures, and the parts it cannot see quietly get worse.

Metered cognition is the arrangement this produces. Reasoning acquires a per-unit price, and the budget sits one or two levels above the person doing the reasoning. It will reach every organisation eventually, because inference costs are real and finance departments are not sentimental. What it moves is the authority to decide how hard a problem is worth thinking about.

Lewis Strauss promised in 1954 that atomic electricity would arrive too cheap to meter. The meter came anyway, and it came first for whoever used the most.

Price something per unit and you have declared it a supply. Supplies get bought from whoever is cheapest, which is exactly the move Microsoft just made. The margin has therefore gone somewhere, and the European earnings published on 5 August say where. SAP's cloud backlog rose 26% at constant currencies to €22.9bn. Capgemini raised its annual guidance on bookings up 9.2%. Sopra Steria upgraded its outlook on 5.3% organic growth. OVHcloud grew public cloud revenue 20.2%. None of these companies builds a frontier model.

What they sell is the part no meter can read: where in a firm carrying thirty years of accumulated process the machine should be connected, which of its outputs a regulator will later ask about, and whose name goes on the decision when it turns out wrong. That work resists a per-token price because nobody buys it by the token. The model was never the scarce part.

The budget itself is a reasonable instrument, and the case for it is stronger than its critics allow. An engineer burning three thousand dollars a month on retries has bought persistence, and persistence is a poor proxy for insight. A company that never measures inference will discover in due course that it has been subsidising noise. Measurement is the correct instinct here. The difficulty is what a token counter can actually see. It sees volume, and it sees nothing about whether the person spending it knew what to ask.

Metered cognition does to a question what budgeting has long done to a role. Jobs leave a company by being defunded, and the work quietly stops being done. No announcement is made that a class of problem is no longer worth investigating. The allocation simply runs dry in the third week, and the cheap answer ships.

Within a year your access to the good model will be a line somebody else approves. The argument that wins you that line is that you know where the machine belongs: which of your firm's thirty years of process it should touch, which of its answers somebody will have to defend, and what happens when it is wrong. SAP and Capgemini are being paid for that judgment at industrial scale. It is available at the level of a single desk too, and it is the only leverage in this whole arrangement that does not get cheaper every quarter.

Thinking now has a unit price. Judgment never had one.