OpenRouter July 2026 LLM leaderboard showing Xiaomi Mimo V2.5 daily volume lead and Chinese model market share near 46%

OpenRouter Rankings July 2026: Who's Actually Winning the AI Model Race

If you still pick an LLM from a benchmark chart you saw in May, you are already behind. OpenRouter publishes what vendors rarely volunteer: real paid token volume across hundreds of models. Through July 25, 2026, Xiaomi Mimo V2.5 leads daily traffic at 1.4T tokens, Chinese labs hold roughly 46% of identified volume, and Hermes Agent alone routes about 45% of app-layer traffic. This guide reads that board as billing truth—not marketing— with model, vendor, pricing, and app tables, a usage-versus-quality barbell analysis, an August outlook, and a five-step tiered routing playbook for OpenClaw teams.

1. July headline: Mimo V2.5, DeepSeek, Hy3, and 46% Chinese share

OpenRouter aggregates production calls from millions of developers. The board reflects what code actually routes in the wild—not a one-off benchmark run. Data is current through July 25, 2026; live rankings shift daily at openrouter.ai/rankings.

The daily token podium as of July 25: Xiaomi Mimo V2.5 at 1.4 trillion tokens/day, DeepSeek V4 Flash at 943.9 billion, and Tencent Hy3 at 590 billion. Seven of the top twelve models are Chinese-origin; NVIDIA Nemotron 3 Ultra (free tier), Claude, and Gemini still anchor the US side.

At provider level, Chinese labs—DeepSeek, Xiaomi, Tencent, Z.ai, MiniMax, Moonshot AI, and Alibaba—now hold roughly 46% of identified token volume, up from under 2% a year ago. US-origin models (OpenAI, Anthropic, Google combined) have fallen from about 70% in mid-2025 to roughly 30–36% in July 2026.

This is pricing math, not narrative. DeepSeek V4 Flash lists near $0.05–$0.14 per million input tokens; GPT-5.5 sits around $5.00—a gap near 35×. When open-weight models are good enough for agent loops and cost a fraction of closed frontiers, developers vote with API keys.

One nuance: DeepSeek remains the most stable #1 provider at roughly 16–18% share, but the "model of the month" crown keeps rotating—MiniMax M2.5 in February, MiMo-V2-Pro in April, Mimo V2.5 in July. Do not treat a single day's #1 as a durable lead.

2. Model Top 12 by daily token volume

Rank Model Vendor Daily tokens 30-day cumulative
1Mimo V2.5Xiaomi1.4T31.2T
2DeepSeek V4 FlashDeepSeek943.9B23.6T
3Hy3Tencent590B23.4T
4Nemotron 3 Ultra 550B (free)NVIDIA428.6B9T
5DeepSeek V4 ProDeepSeek413.7B11.6T
6GLM 5.2Z.ai316.7B13.3T
7MiniMax M3MiniMax262.5B15.1T
8Step 3.7 FlashStepFun204.8B5.9T
9Kimi K3Moonshot AI157.6B1.6T (fastest riser)
10Ling 3.0 FlashInclusionAI (Ant)128.3B417.3B
11Gemini 3 Flash PreviewGoogle106.3B4T
12Claude Sonnet 5Anthropic99.5B3.6T

Rankings move quickly. Claude Opus 4.8 sat in the top ten on July 24 before Ling 3.0 Flash displaced it by July 25—another reason to snapshot weekly rather than cite a single day as gospel.

3. Vendor share: China vs United States

Provider-level shares vary slightly by measurement window and whether first-party app traffic is included. The table below synthesizes multiple seven-day sources into rounded ranges.

Vendor Origin Token share (approx.)
DeepSeekChina16%–18%
XiaomiChina8%–18% (Mimo V2.5 spike)
AnthropicUnited States10%–15%
TencentChina8%–13%
GoogleUnited States8%–13%
Z.aiChina4%–7%
OpenAIUnited States6%–8%
NVIDIAUnited States~5%
MiniMaxChina4%–8%
Moonshot AIChina3%–4%
Alibaba (Qwen)China1%–4%

Chinese labs combined: ~46%. US labs combined: ~30–36%. That twelve-month swing is among the steepest share migrations in the AI infrastructure layer—and it is visible in invoices, not press releases.

4. Usage is not quality: the barbell market

Every OpenRouter recap should lead with this caveat: token volume is not capability. A cheap model behind a high-traffic consumer app can outrank a frontier model teams reserve for the hardest 10% of work.

OpenRouter's spend-by-task breakdown tells a different story than raw tokens. General chat accounts for 35.7% of spend, agentic workflows 30.4%, code 26.5%, and data work 7.5%. Drill into the hardest category—classification and complex reasoning—and the leaders flip: Claude Sonnet 4.6 and Claude Opus 4.7 tie at 13.5% of spend each, with GPT-5.5 third at 11.6%. The cheap open models dominating volume charts barely register here.

Market segment Typical workloads July 2026 leaders Pricing posture
High-volume, error-tolerant Chat, creative writing, roleplay, routine coding assist Mimo V2.5, DeepSeek V4 Flash, Hy3, Nemotron free tier Penny-per-million input; volume wins
Low-volume, low-error-tolerance Classification, complex reasoning, enterprise agent planning Claude Sonnet 4.6/Opus 4.7, GPT-5.5, Claude Opus 5 Dollar-per-million; benchmark-backed

Anthropic's July 24 launch of Claude Opus 5 underscores the premium side: 43.3% on FrontierBench v0.1 versus GPT-5.6 Sol's 37.5%, with Opus-tier pricing at $5/$25 per million tokens (half of Fable 5's input rate). That is a "we cost more and we are worth it" bet—and for hard tasks, benchmarks still back it.

5. App layer: Hermes, Kilo Code, and the hidden roleplay market

Model rankings show which brain is popular. The app leaderboard at openrouter.ai/apps shows what that brain actually does.

Rank App Type Share (approx.)
1Hermes AgentPersonal / CLI agent (Nous Research)~45%
2Kilo CodeCoding agent~13%
3OpenClawGeneral agent~9%
4Claude CodeCoding agent~6%
5DescriptContent production~4.5%
6piAgent~3.3%
7LemonadeCompanion / gaming~2.1%
8ISEKAI ZERORoleplay~2.0%
9Janitor AIRoleplay~1.8%
10ClineCoding agent (IDE)~1.7%

Hermes Agent at ~45% is not a rounding error—it is the single largest app on the platform. Coding agents collectively dominate the rest: Kilo Code, OpenClaw, Claude Code, pi, and Cline all rank in the top ten.

The fork chain Cline → Roo Code → Kilo Code shares one open-source lineage; the youngest fork, Kilo Code, has now overtaken both ancestors in volume. First-mover advantage in open-source dev tooling is thinner than it looks.

The segment enterprise reporting almost never covers: roleplay and companion apps—Janitor AI, ISEKAI ZERO, SillyTavern, HammerAI—move serious volume. OpenRouter's State of AI research with a16z found creative roleplay accounts for more than half of open-model usage on the platform. If your view of AI demand comes only from enterprise headlines, you are missing half the market—and misreading why certain cheap models spike on the volume board.

6. Pricing and positioning table

List prices shift with cache tiers and peak windows. Use this table for routing decisions; verify live numbers before locking budgets. For API setup and SDK examples, see our OpenRouter API tutorial.

Model Input / M tokens Output / M tokens Context Positioning
DeepSeek V4 Flash~$0.05–0.14~$0.24–0.281MCost king; agentic coding default
Nemotron 3 Ultra$0.42 (free tier available)$2.61US open weights; NVIDIA ecosystem
MiniMax M3$0.10$1.21LongMultimodal / image input on a budget
GLM 5.2$0.45$3.31Closest open-weight Opus-style planning
Kimi K3~$3~$151MLargest open weights (1.4TB); frontier-class
Claude Opus 5$5 ($10 fast tier)$25 ($50 fast tier)1MClosed flagship; July benchmark leader

7. August 2026 outlook: five signals to watch

  1. Chinese open-weight combined share likely climbs toward or past 50% unless a major US provider makes a real pricing move. So far, neither OpenAI nor Google has signaled a race to the bottom.
  2. The "model of the month" title keeps rotating. Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot are shipping and re-pricing fast enough that a new daily #1 in August would not be surprising.
  3. Anthropic may ship a cheaper volume tier rather than relying on Opus 5 alone. Opus 5 is Anthropic's fourth flagship in under two months (after Mythos 5, Fable 5, and Sonnet 5)—a cadence aimed at full price-tier coverage, not a single benchmark win.
  4. Kimi K3's 1.4TB open weights will likely see community quantization within two to four weeks, following prior mega-releases. Until then, practical beneficiaries are large teams and inference hosts—not individual laptops.
  5. Security and governance become formal selection criteria. OpenAI's disclosed sandbox-escape incident during internal testing, the proposed US "AI Kill Switch Act," and an expected White House pre-release review framework before August 1 push "vendor safety track record" onto enterprise scorecards—favoring labs with clean histories and pushing tighter permission models for autonomous agents.

8. Three selection pitfalls for platform teams

Pain point 1: Treating volume rank as quality rank. Leaderboard position reflects price sensitivity and app wiring, not fitness for your workload. A model at #1 for roleplay traffic may fail your classification SLA. Build an internal eval before promoting any slug to primary.

Pain point 2: Single-model routing in a rotating market. The July #1 changed between July 24 and July 25. Hard-coding one vendor per quarter guarantees rewrite churn when August releases land. Primary/fallback chains and SecretRef-managed keys are cheaper than agent rewrites.

Pain point 3: Optimizing models while the gateway sleeps. Hermes-style agents and Kilo Code loops assume hours of uninterrupted tool calls. A laptop lid or VPS OOM kill looks like "the model got worse" when the channel simply stopped forwarding. Host uptime—not parameter count—often caps OpenRouter strategy. Cross-reference our June 2026 OpenRouter rankings guide and channels probe runbook when chat goes silent despite green probes.

9. Practical guidance by role

Independent developers and small teams

OpenRouter remains the fastest way to A/B test dozens of models behind one API key—ideal for prototyping. Expect regional latency around 180–250 ms from some geographies, and plan compliance separately for production. For coding, start with DeepSeek V4 Flash for cost and GLM 5.2 for Opus-like planning; reserve Claude Opus 5 or GPT-5.6 for steps where cheaper models actually fail.

Enterprise platform leads

Do not select models by usage rank alone. The barbell in July's data—cheap models for high-volume/low-risk work, premium models for hard/high-stakes work—is a directly actionable routing strategy. Add vendor safety records to your scorecard after this week's governance headlines.

Agent and coding tool builders

The Cline → Roo Code → Kilo Code lineage proves moats in open-source dev tooling are thin. If your product touches entertainment or companion use cases, do not underestimate that segment—the volume is real even when enterprise press ignores it.

10. Five-step tiered routing HowTo

  1. Archive the July baseline. Snapshot model Top 12, vendor share (~46% Chinese / 30–36% US), and app splits (Hermes ~45%, Kilo Code ~13%). Maintain a weekly delta sheet against openrouter.ai/rankings and our weekly token rankings guide.
  2. Map workloads to the barbell. Agent batch jobs and routine codegen → DeepSeek V4 Flash or Mimo V2.5; classification and complex reasoning → Claude Opus 5 or GPT-5.6; ultra-long documents → Kimi K3 or Kimi K2.6; Google-native multimodal → Gemini 3 Flash Preview.
  3. Write primary and fallback chains in openclaw.json. OpenRouter model IDs use vendor prefixes; store keys via SecretRef; enable automatic fallback on HTTP 429 per the channels probe and 429 runbook.
  4. Validate on your own eval set. Measure acceptance rate, error rate, and p95 latency on real prompts—not leaderboard rank—before promoting a model to primary.
  5. Deploy an always-on remote Mac gateway. Run openclaw gateway install under launchd; sync workspaces with SFTP or rsync; pass openclaw channels status --probe weekly as August releases land.
# Example OpenClaw primary / fallback sketch (adjust model IDs to your eval)
openclaw config set models.primary "openrouter/deepseek/deepseek-chat"
openclaw config set models.fallback '["openrouter/xiaomi/mimo-v2.5","openrouter/anthropic/claude-opus-5"]'
openclaw gateway restart
openclaw channels status --probe

11. Frequently asked questions

What was the most-used model on OpenRouter in July 2026? By daily token volume as of July 25, Xiaomi Mimo V2.5 leads at roughly 1.4T tokens per day, followed by DeepSeek V4 Flash (943.9B) and Tencent Hy3 (590B).

Does a high OpenRouter ranking mean a model is the best? No. Rankings measure paid token volume, not capability. Check spend-by-category and your own eval before routing production traffic.

Why do Chinese models dominate volume but US models still win on hard tasks? The market has barbelled: Chinese open-weight models absorb high-volume, error-tolerant workloads at far lower pricing, while Claude and GPT families still lead spend on classification and complex reasoning.

Which apps drive the most OpenRouter tokens? Hermes Agent leads at roughly 45%. Kilo Code (~13%), OpenClaw (~9%), and Claude Code (~6%) follow. Roleplay apps form a large segment rarely covered in enterprise reporting.

12. Conclusion: capability vs popularity

The July story is not "Chinese models won." It is that capability and popularity are diverging. Chinese open-weight labs bought roughly half the platform's volume with price. US closed-frontier labs defend the other side with pricing power on hard tasks, benchmark leadership, and— increasingly—safety credibility.

August will sharpen that split. GPT-6, cheaper Anthropic tiers, Kimi K3 community quantizations, and governance headlines will all land inside weeks, not quarters. The durable skill is not picking this week's #1 model. It is building architecture that switches models without rewriting agents—tiered routing, weekly probes, and eval sets you own.

That architecture still has a physical bottleneck: gateway uptime. Intermittent laptops, sleeping Windows hosts, and memory-starved VPS instances waste leaderboard strategy before the first invoice closes. Rankings tell you which models to route; they do not keep Hermes loops or OpenClaw channels alive at 3 a.m.

SFTPMAC remote Mac rental targets 7×24 OpenRouter and OpenClaw agent workloads on Apple Silicon: native launchd supervision, low-latency API callbacks, and SFTP/rsync baselines aligned with our gateway restart and channel-probe runbooks—a better production home for July's routing strategy than a household machine doubling as an AI gateway.