Author: Claude, Deep潮 TechFlow
DeepCha overview: Data from the AI routing platform OpenRouter shows that the share of tokens used by U.S. companies for Chinese AI models surged from under 5% at the beginning of 2025 to 46% in April 2026, remaining above 30% weekly since February 8. DeepSeek, with a 17.6% market share, has become the platform’s largest single supplier, surpassing Google, Anthropic, and OpenAI. Price is the key driver: DeepSeek V4 Flash costs just $0.14 per million tokens—less than one-thirty-sixth the cost of GPT-5.5. AI startup Lindy has switched 100% of its traffic from Claude to DeepSeek, reducing inference costs by 90%.

One and a half years ago, U.S. companies barely used Chinese AI models. Now, nearly half of the query volume is directed to China.
According to CNBC on July 7, data from the AI routing platform OpenRouter shows that the share of tokens used by U.S. companies to access Chinese AI models on the platform surged from an average of 4.5% in the first half of 2025 to a peak of 46% in April 2026.
Since February 8, this ratio has remained above 30% each week. Meanwhile, the share of the U.S. model dropped from approximately 70% in June 2025 to around 30% in June 2026.
This is not a small-scale experiment by developers. June data from Ramp, an enterprise expense management platform, shows that DeepSeek has topped the list of “trending software vendors,” with U.S. companies directly paying DeepSeek to send data to its API services. On July 1, Palantir CEO Alex Karp publicly criticized the pricing models of U.S. AI labs in a CNBC interview, stating that enterprise customers are “paying for tokens that generate no value.”
Price gap: 60% to 90% cheaper, with a maximum difference of 36 times
Price is the core variable driving this migration.
According to Justin Summerville, a data analyst at OpenRouter, Chinese open-source models are 60% to 90% cheaper than Anthropic and OpenAI’s top products. In terms of pricing, DeepSeek V4 Flash charges just $0.14 per million input tokens, compared to $5 for GPT-5.5—a difference of approximately 36 times. The gap is even wider on the output side: DeepSeek V4 Flash charges $0.28 per million output tokens, while GPT-5.5 charges $30—a difference of more than 100 times.
According to VentureBeat, even DeepSeek’s flagship V4-Pro, at $1.74 per million input tokens, costs roughly one-seventh of GPT-5.5 and about one-sixth of Claude Opus 4.7. With caching enabled, the gap widens further, reducing DeepSeek V4-Pro’s cost to as low as one-tenth of GPT-5.5.

Harpreet Arora, Head of Infrastructure at Vercel AI, told CNBC that after Zhipu AI’s GLM 5.2 was released in June, it achieved the fastest adoption rate on the Vercel platform in 2026, with daily token usage increasing by approximately 27 times and the number of users growing by about 80 times in its first week. Arora’s assessment was straightforward: “Price is working. When tasks don’t require the best model, teams start routing them to the cheapest one that’s good enough.”
Lindy has discontinued Claude and fully switched to DeepSeek, reducing inference costs by 90%.
The AI startup Lindy is the most representative case in this migration.
This 25-person AI agent company previously relied entirely on Anthropic’s Claude model. CEO Flo Crivello announced on X that the company has migrated 100% of its traffic to DeepSeek v4, hosted within the United States by U.S.-based provider Atlas Cloud. Crivello told CNBC that after the switch, “the cost curve plummeted,” saving the company millions of dollars and reducing inference costs by approximately 90%.
According to The New Stack, Lindy’s previous AI inference costs had exceeded personnel expenses, and Crivello said this was “a matter of survival” for the company. He stated that if Anthropic lowers its prices, he would be willing to switch back. But until then, the company has no other option.
Lindy is not an isolated case. According to CNBC, Uber burned through its entire annual AI budget in just four months in 2026, primarily due to Claude Code. GitHub also faced runaway costs from Copilot’s agent mode and was forced to shift from a flat monthly fee to a pay-as-you-go model.
DeepSeek tops Ramp's enterprise spending leaderboard, transitioning from "trial" to "purchase"
OpenRouter shows token flows at the developer level. Ramp’s data reveals a more significant signal: Chinese AI models are entering the formal procurement processes of U.S. enterprises.
According to Ramp’s June report, DeepSeek has risen to number one on the “Trending Software Vendors” ranking, which is based on real transaction data from over 50,000 U.S. companies and measures explosive growth in first-time purchases. Ramp’s Chief Economist, Ara Kharazian, noted that U.S. companies are no longer merely downloading DeepSeek’s open-source models for self-deployment—they are now directly paying DeepSeek to send and receive data via its API.
Kharazian views cost awareness as the primary catalyst for this wave of adoption. DeepSeek briefly reached a 0.3% enterprise penetration rate in January 2025 with the release of R1, before declining to 0.1%. The current resurgence is driven by more substantial factors: in May, DeepSeek made permanent discounts on its V4-Pro model, reducing cached input pricing to approximately $0.0035 per million tokens.

The U.S. model's market share has halved in a year, as the market splits into a "commodity layer" and a "premium layer."
From a platform-wide perspective, the speed of this share migration is astonishing.
According to OfficeChai, citing OpenRouter data, U.S. models (combined Google, OpenAI, and Anthropic) accounted for approximately 70% of token usage on OpenRouter in June 2025. By June 2026, this share had dropped to approximately 30%. DeepSeek became the platform’s largest single provider with a 17.6% token share, followed by Alibaba’s Qwen at 13.9%. Chinese models collectively accounted for approximately 44% of total token traffic among the top ten models.
OpenRouter's own scale is also rapidly expanding.
According to data cited by Bloomberg, the platform’s weekly token processing volume increased fourfold, from approximately $5 trillion in April 2025 to over $20 trillion in April 2026. The share of programming workloads surged from 11% at the beginning of 2025 to over 50% by mid-2026, with Chinese models demonstrating particular cost-effectiveness on programming tasks.
However, token share does not equate to revenue share. Although Anthropic’s Claude has been squeezed to approximately 13% in token share, its pricing per token is significantly higher than that of Chinese open-source models, resulting in a revenue share far exceeding what its token share suggests. The market is splitting into two tiers: the premium tier, dominated by U.S. closed-source models that monetize through capability premiums; and the commodity tier, led by Chinese open-source models that compete on price and scale.
The enterprise AI cost crisis spreads as Palantir’s CEO publicly criticizes token pricing
Cost pressures have spread from startups to large enterprises.
Palantir CEO Alex Karp publicly criticized OpenAI and Anthropic’s token pricing model on CNBC’s Squawk Box on July 1. Karp stated that U.S. companies are paying for “tokens that generate no value,” and that their intellectual property and competitive advantages are flowing to AI labs. The day before the interview, Palantir released nine “AI Sovereignty” declarations, condemning “tokenmaxxing”—the excessive consumption of tokens to maximize AI usage—as delivering only “false progress.”
Behind Karp’s remarks are real business pain points. As AI workflows shift from simple conversations to an “agent” model—where models autonomously plan, invoke tools, and execute multi-step tasks—the token consumption per task has increased 10 to 30 times. OpenAI CEO Sam Altman recently acknowledged that AI costs have become a “major issue” for enterprise customers.
The Linux Foundation established the Tokenomics Foundation, backed by companies such as Google, Microsoft, IBM, and Salesforce, to create an open standard for AI token costs. This alone demonstrates that enterprises currently lack a unified method to measure AI expenditures.
For U.S. companies, the outcome is ironic: the government’s attempt to curb China’s AI development may be driving its own business customers toward Chinese models.
