LLM Rankings
50 models · 67.07T tokens this week · ranked by real usage
Key facts · week of 2026-08-10
The most-used LLM for the week of is DeepSeek: DeepSeek V4 Flash 0731, with 15.8% of OpenRouter token traffic (10.69T tokens).
§The top 3 models (DeepSeek: DeepSeek V4 Flash 0731, Tencent: Hy3, OpenAI: GPT-5.6 Luna (batch)) account for 39.2% of all usage.
§This week's biggest riser is OpenAI: GPT-5.5 (batch), up from #50 to #39.
§The cheapest model in the usage top 10 is OpenAI: GPT-5.6 Luna (batch) at $0.09999999999999999 per 1M input tokens (usage rank #3).
§The largest context window in the top 10 is OpenAI: GPT-5.6 Luna (batch) at 1.1M tokens.
§
Cite as: AI Tier List, LLM Usage Rankings (2026-08-10), www.aitierlist.xyz/en/models — data: OpenRouter
Which model should you pick?
- When quality comes first
- Anthropic: Claude Fable 5 (batch)
- Artificial Analysis intelligence index score 62.1 — highest in this ranking
- The proven default
- DeepSeek: DeepSeek V4 Flash 0731
- 15.8% usage share — #1 in real usage
- When cost matters
- NVIDIA: Nemotron 3 Ultra (free)
- Cheapest in the top 15 — Free/1M input
- For long docs & codebases
- OpenAI: GPT-5.6 Luna (batch)
- 1.1M token context window
- To bet on the trending choice
- OpenAI: GPT-5.5 (batch)
- Up 11 spots (#50 → #39)
Full ranking · 50 models
| Rank | Model | Share |
|---|---|---|
1 | DeepSeek: DeepSeek V4 Flash 0731 | |
2 | Tencent: Hy3 | |
3 | OpenAI: GPT-5.6 Luna (batch) | |
4 | DeepSeek: DeepSeek V4 Flash 0423 | |
5 | Xiaomi: MiMo-V2.5 | |
6 | Z.ai: GLM 5.2 (batch) | |
7 | DeepSeek: DeepSeek V4 Pro | |
8 | NVIDIA: Nemotron 3 Ultra (free) | |
9 | Google: Gemini 3.6 Flash (batch) | |
10 | Poolside: Laguna S 2.1 (free) | |
11 | Claude Opus 5 (batch) | |
12 | MiniMax: MiniMax M3 (batch) | |
13 | MoonshotAI: Kimi K3 | |
14 | StepFun: Step 3.7 Flash | |
15 | Anthropic: Claude Sonnet 5 (batch) | |
16 | Google: Gemini 3 Flash Preview (batch) | |
17 | OpenAI: GPT-5.6 Terra (batch) | |
18 | Anthropic: Claude Sonnet 4.6 (batch) | |
19 | OpenAI: GPT-5.6 Sol (batch) | |
20 | Google: Gemini 2.5 Flash Lite (batch) | |
21 | Anthropic: Claude Opus 4.8 (batch) | |
22 | OpenAI: GPT-5.6 Luna Pro (batch) | |
23 | Google: Gemma 4 31B (free) | |
24 | Google: Gemini 2.5 Flash (batch) | |
25 | Xiaomi: MiMo-V2.5-Pro | |
26 | OpenAI: gpt-oss-120b | |
27 | Google: Gemini 3.1 Flash Lite (batch) | |
28 | DeepSeek: DeepSeek V3.2 | |
29 | Google: Gemma 4 26B A4B (free) | |
30 | NVIDIA: Nemotron 3 Super (free) | |
31 | SpaceXAI: Grok 4.5 | |
32 | Anthropic: Claude Opus 4.7 (batch) | |
33 | Qwen: Qwen3.8 Max | |
34 | NVIDIA: Nemotron 3.5 Lightning (free) | |
35 | Anthropic: Claude Fable 5 (batch) | |
36 | Anthropic: Claude Haiku 4.5 (batch) | |
37 | inclusionAI: Ling-2.6-flash | |
38 | Cohere: North Mini Code (free) | |
39 | OpenAI: GPT-5.5 (batch) | |
40 | OpenAI: GPT-4o-mini (batch) | |
41 | Google: Gemini 3.5 Flash Lite (batch) | |
42 | MiniMax: MiniMax M2.7 | |
43 | Anthropic: Claude Opus 4.6 (batch) | |
44 | OpenAI: GPT-5 Mini (batch) | |
45 | DeepSeek: DeepSeek V4 Pro 0813 | |
46 | OpenAI: GPT-5.4 (batch) | |
47 | Poolside: Laguna XS 2.1 (free) | |
48 | Google: Gemini 3.1 Pro Preview (batch) | |
49 | OpenAI: gpt-oss-20b (free) | |
50 | Google: Gemini 3.5 Flash (batch) |
Model Comparisons
- DeepSeek: DeepSeek V4 Flash 0731 vs Tencent: Hy3
- DeepSeek: DeepSeek V4 Flash 0731 vs OpenAI: GPT-5.6 Luna (batch)
- DeepSeek: DeepSeek V4 Flash 0731 vs DeepSeek: DeepSeek V4 Flash 0423
- Tencent: Hy3 vs OpenAI: GPT-5.6 Luna (batch)
- Tencent: Hy3 vs DeepSeek: DeepSeek V4 Flash 0423
- OpenAI: GPT-5.6 Luna (batch) vs DeepSeek: DeepSeek V4 Flash 0423
- Full LLM API pricing →
Related AI Tool Tier Lists
Get next week’s LLM rankings
These rankings refresh every week. We email you when they move — once a week, no spam.
FAQ
What is the most-used LLM?
DeepSeek: DeepSeek V4 Flash 0731 by deepseek ranks first by real token usage. It accounts for roughly 15.8% of tokens in this segment. Tencent: Hy3, OpenAI: GPT-5.6 Luna (batch), DeepSeek: DeepSeek V4 Flash 0423 follow.
How are these rankings calculated?
Models are ranked by tokens actually consumed through OpenRouter, not by benchmark scores. It refreshes weekly and currently tracks 50 models.
Which is the cheapest LLM?
Among the paid models in this ranking, inclusionAI: Ling-2.6-flash is cheapest at $0.01 per 1M input tokens. 9 models in the list are free to use.
Which model has the longest context window?
OpenAI: GPT-5.6 Luna (batch) has the longest context window in this ranking at 1.1M tokens.
Token usage and share data aggregated by OpenRouter, refreshed weekly. Prices are USD per 1M tokens. Benchmark scores come from Artificial Analysis; models it hasn't measured show —.