Gemini has over a billion users worldwide. This service processes 22 billion tokens every minute.
However, in recent months, the company that created this service has received a strange evaluation: It is being labeled as falling behind in the AI race.
If you have ever searched for something in a search bar, you are already using this company’s service. The same applies if you use Gmail, an Android phone, or watch YouTube. What this company decides to do next is closer to your daily life than any benchmark leaderboard.
Photo: StockCake (CC0, Free to use)
Evidence of Falling Behind
Evidence of this decline has appeared on three fronts simultaneously.
The first sign was in the product itself. The top-tier model, Gemini 1.5 Pro, was teased at a developer event with a ‘coming next month’ promise. That month passed, and several more followed. In the meantime, mostly lightweight and fast ‘Flash’ models were released.
The second sign came from developers’ wallets. Looking at spending share in the coding tool market, Anthropic’s Claude Code holds 42%, and OpenAI’s Codex holds 21%. Google’s share falls significantly short of these two.
Coding is not just a simple trick. It is the ability to break down problems, call tools independently, and fix errors. Falling behind in this capability is akin to handing over leadership in the next-generation AI assistant market.
The third sign was the people. Noam Shazeer and John Jumper left the company. Jeff Dean, a 27-year veteran, also left to pursue a startup.
Demis Hassabis moved to the position of DeepMind CEO to focus on drug discovery. The vacancy was filled by Koray Kavukcuoglu, VP of Product and Commercialization. While the name DeepMind remains, real power has shifted to other hands.
For someone who spent 27 years at one company to walk away and start over from scratch—this is not just an exceptional departure, but a classic indicator of where the organization’s center of gravity is shifting.
The number of 22 billion tokens per minute is hard to grasp. Divided by seconds, that’s about 360 million. If we assume it takes roughly 150 tokens to fill one page of 200-character manuscript paper, this company is processing the equivalent of 2.4 million pages of text every second.
Why does a company handling this scale seem to keep falling behind in the race for benchmark #1?
The Species That Changes the Rules
In biology, there is a concept called ’niche construction.’ Natural selection is usually explained as a process where the best-adapted individuals survive. However, some species use a different strategy. Instead of adapting to the environment, they change the environment itself.
Beavers don’t compete to be the predator with the sharpest teeth. Instead, they build dams. They alter the flow of the river and create ponds, making those ponds a habitat favorable to them. The winner is not the individual that chews wood best, but the one that changes the geography of the entire river.
The competition currently happening between OpenAI and Anthropic is closer to ‘who has the sharper teeth.’
The #1 spot in benchmarks changes every few weeks, or months at most. If you don’t win this competition, it’s hard to guarantee investment or survival. This is because building the model itself is the entire business.
Google stands in a different place. Search, Android, YouTube, Workspace—it already owns the entire river. What this company needs is not the sharpest teeth, but a dam that makes the river flow in its favor.
This pattern is not unfamiliar in the history of the tech industry. It has been repeated that companies with distribution networks and ecosystems, rather than those that build the most outstanding products, eventually absorb the market. Making the best object and owning the land people walk on every day are different games.
How to Build a Dam
The dam is being built on four fronts simultaneously.
The first point is the redesign of the models themselves. Instead of deploying the largest model for every query, Google breaks down models by task. They set ‘Flash’ models—which outperform the previous generation’s top-tier models while being four times faster and significantly cheaper—as the default. They only pull out the heavy models for truly difficult problems.
Second is chip design. Starting with the 8th generation TPU, they have completely separated ’training’ and ‘inference’ chips. This decision itself is a declaration. It is a declaration that the battleground of AI has shifted from large-scale training that takes months, to the daily, billions-of-times-repeated responses—inference.
Third is the form of the product. The direction represented by ‘Gemini Spark’ goes beyond a chatbot that answers questions. The goal is a 24-hour agent that works organically with Gmail, Drive, Calendar, YouTube, and Chrome to handle practical tasks.
For example: It reads a meeting request in your inbox, finds an empty slot in your calendar, sends a reply, and finds relevant documents in your Drive to organize them in advance. Instead of the user opening and closing chatbots, the service works continuously in the background.
And perhaps the smartest point is the last one. They attach targeted ads to answers in Search AI mode. For queries with clear purchase intent, such as loans or insurance, the probability of an ad appearing exceeds 50%.
To competitors, longer answers mean increased token costs. To Google, longer answers mean more space to attach ads.
This change comes to consumers in two ways. One is that they can use free or cheap AI services longer and more widely. The other is that ads are being blended into those answers in increasingly natural forms.
Photo: StockCake (CC0, Free to use)
However, a dam comes with conditions.
‘Similar performance but cheaper’ is innovation. ‘Inferior performance but cheaper’ is just a low-end model. Google’s entire strategy is built on this one line. If the premise that the Flash model actually gets the job done and the agent actually performs the work falls apart, all that remains is a reputation for being ‘cheap.’
A more fundamental problem exists. Why do top-tier researchers choose their labs? Noam Shazeer, John Jumper, and Jeff Dean all eventually moved to pursue the frontier. Organizations that aim to build the world’s best models attract the world’s best talent. A #2 organization that prioritizes cost and efficiency, at least by the calculations of those talents, lacks a comparable magnet.
The gap that has already opened in the coding agent market points to the same problem. No matter how well-equipped the infrastructure is, if you fall behind in the tools that developers use daily, you end up handing over the agent market that would run on that infrastructure to competitors. Changing the geography of the river and keeping fish alive in that river are different problems.
If we go one step further, we can see why Google had no choice but to make this selection.
Suppose we apply a giant model to all 1 billion users. The computational resources on Earth cannot handle it. Being large is both a freedom and a prison.
OpenAI and Anthropic can afford to push for top performance because they have relatively fewer users to lose. Google is already responsible for too many people, so it had to voluntarily give up that freedom.
1 billion users, 22 billion tokens per minute. When I first saw these numbers, they seemed at odds with the assessment that Google was falling behind.
Looking at them again, they read differently. These are not numbers on the opposite side of failure, but evidence that they are playing a different game entirely.
The problem is that the winner of this game has not been decided yet. The maturity of coding agents and the profitability of the service ecosystem will provide answers over the next few years. Whether the company that builds the smartest model will win, or the company that provides the most affordable, widely accessible, daily-use service will win, is yet to be determined.
Photo: StockCake (CC0, Free to use)
Next time you ask something in a search bar and get a long answer, it’s worth thinking about:
Is this answer long because of an effort to be more accurate, or because they needed one more spot to attach an ad?
References
- Axios report — Coverage on the delay of Google Gemini 1.5 Pro
- Fortune report — Coverage on key AI talent departures at Google
- OpenRouter coding tool spending share data (Claude Code, Codex, etc.)
- Report on Google DeepMind reorganization and transition to Koray Kavukcuoglu leadership
- Google Gemini user metrics — 1 billion MAUs, 22 billion tokens processed per minute
- Announcement on Google 8th Gen TPU training/inference separation architecture
- Materials on Gemini Spark agent product direction
- Materials on Search AI mode ad integration and matching rates
- T-Times TV — Analysis report on Google's changed AI business strategy