AI: OpenAI, Google and Anthropic shift the battle to speed

AI: OpenAI, Google and Anthropic shift the battle to speed

Model generation speed is starting to become a real priority for AI vendors. OpenAI, Anthropic, Google… The publishers’ guideline is no longer just the intelligence of the models.

What if the raw power of the models was no longer the only criterion for the success of a good LLM? With the era of agentic AI, generative model editors are waging a new battle: that of speed. The goal is no longer to deliver the best intelligence at all costs, but rapid intelligence. The goal: to produce more, faster, in all areas, from code to research and scientific discovery, including process automation. A race within the race that all the model editors are taking part in.

Intelligence is commoditized

May 2026, Anthropic releases Claude Opus 4.8, its most powerful model. Price: 5 dollars per million tokens in, 25 in output. Exactly the price of GPT-5.5, OpenAI’s spearhead. Eighteen months ago, a publisher’s high-end model often cost several times that of the neighbor’s. This gap has disappeared. At the top, the prices have aligned, and the benchmark scores too: on reasoning, code or agentics, Opus 4.8, GPT-5.5 and Gemini are within a few points. “We are starting to move towards a common intelligence gap,” confirms Hamidou Dia, VP Applied AI Engineering at Google Cloud. Since the advent of agents, mostly code, users no longer just talk to models, they put them to work in the real world. The model does not respond all at once: it plans, executes, checks, corrects, and restarts as many times as necessary. Each step triggers a call to the model, and each call waits until the previous one has finished. As a result, the latencies add up.

Publishers have understood this, and since the start of the year we have seen a series of announcements around speed. As of January 14, OpenAI unveils an agreement with Cerebras to add 750 MW of “ultra-low latency” computing to its platform. In February, Anthropic launched its fast mode: the same intelligence of Opus, but 2.5 times more tokens per second, at an additional cost. On March 5, it’s the turn of OpenAI Codex to unleash its quick mode. With /fast, GPT-5.4 runs 1.5 times faster, with identical intelligence and reasoning. Finally, on May 19, at I/O, Google released Gemini 3.5 Flash, marketed as “four times faster than other frontier models” in its category while beating its own Gemini 3.1 Pro on agentic benchmarks.

A top priority for engineering teams

At OpenAI, we take priority head-on. “We have been focused on speed for a while,” assures Thibault Sottiaux, Core Product & Platform manager at OpenAI, who works on Codex. The editor goes so far as to recognize an unflattering starting point. “Six months ago, everyone was saying that Codex was slow, unusable. We said okay, let’s go ahead and solve that by going back to the fundamentals,” he confides. And the work is far from being closed: “There are a large number of new techniques that we are in the process of scaling and on which we have not yet communicated.”

Same position at Google where latency is now a “criterion on which the teams are hyper focused”, insists Hamidou Dia. The publisher has also divided its hardware accordingly: “We now have a TPU dedicated to training, a TPU dedicated to inference,” he recalls. The instructions even go down to the development teams. “Within Google, in some of our development teams, we have a latency budget. For example, we ask you to optimize a process in 10 seconds, and if you manage to do it in 5, you have 5 seconds that you can spend elsewhere”, Hamidou Dia.

Same thinking at Anthropic, which is working on the issue of speed on several levels: a range that goes from the Haiku model, light and fast, to Opus and its paid fast mode. Because the demand comes from customers, whose reading grid has changed. “Many users are looking for the best intelligence per dollar per second,” recalls Katelyn Lesse, head of Engineering at Claude Platform. But the issue goes beyond the simple adjustment slider. For certain abilities, speed directly determines usage. “There are areas, notably computer use, where an acceleration compared to current speeds is really necessary,” illustrates Angela Jiang, head of Product at Claude Platform.

It remains to know where to place the cursor. At Nvidia, which equips almost the entire industry, we are moving the priority from the model to the system. “Speed ​​is not just a question of where the agent is turning, it’s also where the tools they want to use are, and how they access them,” explains Nat Ives, director France, Benelux and Nordics at Nvidia. Because the speed of an agent does not depend only on the GPU: once launched, it must use tools, the CPU, the network, storage, request a downstream system, and it is often this link which dictates its speed.

New indicators emerge

The same movement therefore crosses the entire industry: it is no longer the model alone that we optimize, but the entire chain. And the boundary between the teams is moving. Where model development and infrastructure engineering used to move in parallel, they now work hand in hand. At OpenAI as at Anthropic, the teams who design the harness, the software layer, collaborate as closely as possible with those who train the models. Google is pushing the needle even further: its silicon teams design TPUs in close collaboration with those of DeepMind.

But speed is just one metric that publishers are now scrutinizing. Because the challenge for the next few years will not only be to produce quickly, but to produce without burning all of your cash. Dario Amodei, the boss of Anthropic, reminded us recently: if revenue growth slips for even one year compared to the amounts swallowed up in computing, a publisher can find itself on the verge of bankruptcy. In this context, a new ratio is emerging in conversations: intelligence per watt. A new race to come.

Jake Thompson
Jake Thompson
Growing up in Seattle, I've always been intrigued by the ever-evolving digital landscape and its impacts on our world. With a background in computer science and business from MIT, I've spent the last decade working with tech companies and writing about technological advancements. I'm passionate about uncovering how innovation and digitalization are reshaping industries, and I feel privileged to share these insights through MeshedSociety.com.

Leave a Comment