open-models

Open weights win the tokens, Anthropic keeps 64% of spend

Promtime

open-models

Open-weight models did most of the work on Vercel's AI Gateway in August and collected almost none of the money. They handled 56% of all tokens routed through the gateway, their first monthly majority, and 14 cents of every estimated dollar spent, while Anthropic took 64 cents, according to Vercel's September report as covered by The New Stack.

At a glance

  • Vercel's gateway routes tens of trillions of tokens a month, and the open-weight share of them went from 7% in December 2025 to 13% in April and 36% in July.
  • The average price per token across the gateway fell 23.2% in August, a third straight monthly decline; among teams processing over 10 million tokens in both July and August, the median fell 7.6%.
  • Google's share of token volume fell from 30% to 5%, with the decline in Gemini 3 Flash accounting for 22 of those 25 points; more than three-quarters of it moved to rival labs.

If you have not been watching the token counts: the pitch for open weights is that you download the model, run it where you like and keep control of your data, usually for less money. The New Stack reports they now trail the leading frontier models by only four to five months, and reported on Monday that open-weight models took 60% of OpenRouter's US token consumption in August, most of it from Chinese-developed models.

The open-weight share of tokens went from 7% to 56% in nine months

Vercel's own series runs like this: 7% in December 2025, 13% by April, and a rise in every month after that. The report published in August put July at 36%. August crossed the halfway line at 56%.

The peak came mid-month. Vercel CEO Guillermo Rauch posted last month that August 22 had been a "record day for open weight share of tokens on Vercel AI Gateway," accounting for 62% of traffic, and read the milestone as an early indication rather than a ceiling.

This is very likely just the start, because enterprise adoption is still early, and harnesses, CLIs, IDEs, SDKs, etc need to be adapted to be model agnostic.

Anthropic's share of spend has not fallen below 61% since December 2025

In August its models took 64 cents of every dollar that moved through the gateway. Vercel says Anthropic has held the top two positions by spend every month since December 2025, and often the third spot as well.

The churn happened inside that slice. Fable 5 fell from 13.2% of total gateway spend in July to 4.9% in August, while the cheaper Opus 5 climbed to 22.5%. Vercel's data shows 90% of teams using Fable cutting back, with more of that work going to Opus 5 than to any other model, and Opus gaining almost twice as much usage as Fable lost. Vercel attributes that to the newer model handling similar workloads at roughly half the price.

Lab loyalty doesn't follow brand, it follows model profile, and consistency wins.

Anthropic priced the swap itself: Opus 5 at half the cost per task

The shift Vercel measured matches what Anthropic said when it launched the model. Anthropic's own announcement of Claude Opus 5 says it comes close to the frontier intelligence of Claude Fable 5 at half the price, and describes Opus 5 as giving greatly improved performance for the same cost as its predecessor, Opus 4.8. On CursorBench 3.2, at max effort, the same announcement puts Opus 5 within 0.5% of Fable 5's peak score at half the cost per task. So the dollars stayed with the lab while customers walked down its own price list.

Google's token volume on the gateway fell from 30% to 5%

Nearly all of that came from one model: the decline in Gemini 3 Flash alone accounted for 22 of the 25 percentage points. More than three-quarters of the volume Gemini 3 Flash lost went to models from other providers, including OpenAI and Anthropic. Vercel's authors put the rule plainly: when a new model preserves what users valued in its predecessor, the lab keeps its customers; when it does not, those customers fill the need elsewhere. The other side of the same pattern showed up at Z.ai, where GLM-5.3-Flash was processing three times the daily volume of GLM-5.2 within five days of launch.

Vercel counts reasoning and cached tokens in the same total

Vercel launched AI Gateway last year so developers could reach models from several providers through one interface, without juggling separate API keys, accounts and rate limits. It sits between the application and the model provider, routes each request, and tracks usage and cost along the way, which is why it can publish splits like these at all. Token volume here counts input and output tokens plus reasoning, cached-input and cache-creation tokens. It is a good proxy for how much work each model is being handed, and a bad one for money: it is a toll booth counting vehicles, not fares, and the open-weight models from DeepSeek, Moonshot AI and Z.ai are the cheap vehicles. The New Stack notes that routers exist because applications have traditionally hard-coded one model for everything, while a router can pick per request.

What the report does not give you is a denominator. Vercel publishes shares of an estimated dollar and a monthly volume of "tens of trillions of tokens," never the actual spend, and the view covers applications running on its own infrastructure rather than the market. In our view the more telling pair is the two price declines: a 23.2% drop in the average price per token next to a 7.6% median drop for teams above 10 million tokens suggests the headline number is moving with the mix of models, not with what individual teams pay.

When the October numbers land

The next monthly report covers September, and three lines are worth watching: whether open weights push past 14 cents of the dollar, whether Anthropic holds its floor of 61% of spend, and whether Google's 5% of token volume recovers. The plumbing is moving too. The New Stack reports that Stripe has announced plans to acquire OpenRouter in a reported $8 billion deal; that one is agreed, not closed.

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.