Sundar Pichai, the CEO of Google, said this on stage in May 2026, during the I/O conference, and he said it as a joke, but you could feel the laugh was a little nervous: "a lot of companies have already blown through their annual token budget, and it's only May." And well that wasn't really a joke. Uber burned through its AI budget for all of 2026 in four months. ServiceNow did the same. And it's not an isolated accident... it's the economic model of closed AI starting to show, and it's showing more and more.

The problem is that the providers have changed the way they bill. For months, they "subsidized" usage to gain market share, as Gabriel Hubert, the cofounder of Dust, puts it so well... and well now that demand is exploding, the subsidy is ending. Etienne Grass, at Capgemini, summed it up in one sentence that sticks: "we're leaving the era of the free lunch." And that's only the beginning, he warned.

500€
Per month, per user
For an AI assistant in heavy cloud usage. 2,000€ for code work. Source: Vincent Luciani, Artefact.

When I see the quotes come through, I tell myself something isn't right. Vincent Luciani, the cofounder of Artefact, put a number on it very precisely: 500 euros per month per user for an AI assistant in heavy usage, and 2,000 euros for code work. So do the math for your company... 50 users at 500 euros, that's 25,000 euros per month, so 300,000 euros per year. And that's with today's prices, because it's going to keep climbing.

On the other side, an on-premise server with a decent GPU costs 20,000 euros, once, amortized over three years. After that, it's just electricity. The gap is an order of magnitude, not just a margin. And in the second year, local costs zero, while the cloud keeps running.

0€
Marginal cost of a local agent
100 agent steps on your server = 0 tokens billed. The model runs on your machine. No API, no meter.

Token maxxing, or when adoption becomes absurd

A few months ago, Silicon Valley invented a word that I find pretty revealing: "token maxxing." It was the trend, in tech companies, to maximize the number of tokens consumed per employee to show that you were using a lot of AI. Some leadership teams even posted internal dashboards with consumption scores per person, and employees were rewarded if they consumed more. It was considered a positive sign of adoption.

And well Bank of America just declared, late June 2026: "it's the end of token maxxing." Amazon stopped displaying token scores per employee. Salesforce launched a tool to measure activity by task completed rather than by token volume. A BCG executive acknowledged a share of "waste" in companies' use of AI... waste, the exact word.

And this is where I want us to think for a second: the economic model of closed AI is a model where the incentive is toward inflation. The provider wants you to consume more. Your employees are rewarded if they consume more. Everyone has an interest in the bill going up, except you. In a local model, there is no token maxxing, simply because there are no tokens... the incentive is toward efficiency, not consumption.

Agents: the multiplier that changes everything

If you think 500 euros per user is already a lot, wait until agents arrive. A chatbot is one question and one answer, so it's controllable. An agent is different: it chains dozens, hundreds of queries to accomplish a task, whether that's coding an app, writing a report, or analyzing a dataset... and each step is billed.

Google now processes 3.2 million billion tokens per month, that's seven times more than a year ago, and three hundred and thirty times more than two years ago. At OpenAI, daily interactions with Codex have multiplied by 15 since January 2026. 179 enterprise customers have crossed the 1,000 billion tokens generated in total... and we're only at the first deployments of agents in business.

On a cloud, each step of the agent is a billed query: 100 steps is 100 times the cost of a chatbot. On a local server, an agent that runs 100 steps costs nothing more than electricity, because the model runs on your machine, each iteration is local computation, no API, no meter. The more the agent iterates, the more the gap widens between cloud and local, it's mathematical.

The Atlanta Fed confirms the trend

In March 2026, the Atlanta Fed published a survey that anticipates a rise in the average AI cost per employee of more than 50% in 2026, to 1,776 dollars per year, in companies with more than 250 employees in the United States. And that's the average... the heavy consumers pull the average up, but the average itself will keep rising because agents are going to become widespread.

Arthur Mensch, the head of Mistral AI, said something in front of lawmakers that I find pretty clear: there is a "risk of lock-in" in the market, with "digital operators raising their prices extremely hard and an inability of players to free themselves from these new services." Read that last sentence carefully... "an inability to free themselves." That's a current state of affairs, and really not a future risk!

The alternative is here, it's technical, and it works

So here's where I stand. Local AI is an advance, not a strategic retreat, and that changes everything... cloud costs are going to explode, not come down, open source models are going to multiply, not close up, and hardware gets more capable every quarter. An open source model served by Ollama, running on a Mac or on a server in your company: fixed cost, unlimited usage, data that stays with you. It's here, it's technical, and it works.

I see this in the companies that reach out to me: the cloud bill becomes a leadership topic, and local AI becomes a concrete option, not a curiosity anymore. It's an economic logic that's hard to argue with, and if you want to talk about it for your situation, I'm at your disposal.