GLM-5.3 vs Kimi K3 vs Qwen3.8 Max vs DeepSeek V4
The four flagships behind almost every cheap coding pass. Two have genuinely open weights, one has an announcement, and one is unconfirmed.
Updated August 28, 2026 · Independent comparison — GLM-5.3, Kimi K3, Qwen3.8 Max, DeepSeek V4 are separate products.
GLM-5.3
Z.ai's flagship, launched 2026-08-14
GLM-5.3 is Z.ai's flagship, live on the API at $1.40 in and $4.40 out per million tokens — the same price as GLM-5.2 — with cache reads at $0.26 and cache storage free for a limited time. Z.ai does not publish an official context figure for 5.3. It reaches developers through the Z.ai API, the GLM Coding Plan (which auto-routes 5.2 and 5.1 requests up to 5.3), ClinePass and OpenCode Go, where the cheaper GLM-5.3-Flash runs at $0.15/$0.50.
Kimi K3
Moonshot's 1M-context flagship
Kimi K3 is Moonshot's flagship at $3 in and $15 out per million tokens, cache reads at $0.30, with a 1M-token context. It is reachable through the Moonshot API, the Kimi membership from the Moderato tier (though the full 1M chat context is gated to the top tier), ClinePass, OpenCode Go and Zen, and Chutes in confidential compute at the full 1M. A weight release has been announced but has no date, so K3 is not yet an open-weight model.
Qwen3.8 Max
Alibaba's flagship, released August 2026
Qwen3.8 Max was released around 2026-08-03 at $2 in and $6 out per million tokens, with cache reads at $0.25 and explicit cache reads at $0.17. Context is 1M with roughly 128K output. It is served from QwenCloud and DashScope on both OpenAI- and Anthropic-compatible endpoints, and appears in ClinePass, OpenCode Go and the QwenCloud Token Plan, where night-time usage between 22:00 and 08:00 UTC+8 burns credits at half rate. Open weights were claimed shortly after launch but remain unconfirmed.
DeepSeek V4
1M context, 384K output, peak/off-peak pricing
DeepSeek V4 is the most aggressively priced flagship here, and the only one that charges differently by time of day. Flash costs $0.22 in and $0.66 out per million tokens off-peak, doubling during peak hours (Mon-Fri 01:00-04:00 and 06:00-10:00 UTC); Pro runs $0.66/$1.98 off-peak and $1.32/$3.96 at peak. Cache hits are close to free at $0.007-$0.044. It offers a 1M context with up to 384K output — the largest output ceiling in this group — across OpenAI-, Anthropic- and Responses-compatible endpoints.
Bottom line
These four models sit behind almost every cheap coding pass, and they are not equally open. GLM-5.3 is the cheapest credible option at $1.40/$4.40 with the widest agent support, though Z.ai does not publish its context window. DeepSeek V4 is cheaper still and uniquely time-of-day priced — Flash from $0.22/$0.66 off-peak — with the largest output ceiling here at 384K on a 1M context, and near-free cache hits. Qwen3.8 Max is $2/$6 with a 1M context and about 128K output; open weights were claimed shortly after its August launch but remain unconfirmed. Kimi K3 is the most expensive at $3/$15, offers 1M context, and despite an announced weight release has no date — so it is not an open-weight model today.
Pricing
| Feature | GLM-5.3 | Kimi K3 | Qwen3.8 Max | DeepSeek V4 |
|---|---|---|---|---|
| Price per Mtok (in / out) | $1.40 in / $4.40 out per Mtok | $3 in / $15 out per Mtok | $2 in / $6 out per Mtok | Flash $0.22/$0.66 off-peak, $0.44/$1.32 peak; Pro $0.66/$1.98 off-peak, $1.32/$3.96 peak |
| Cache pricing | $0.26 read; storage free for a limited time | $0.30 read | $0.25 read; $0.17 explicit cache read | $0.007-$0.044 per Mtok — effectively free |
Capacity
| Feature | GLM-5.3 | Kimi K3 | Qwen3.8 Max | DeepSeek V4 |
|---|---|---|---|---|
| Context window | Not published | 1M | 1M | 1M |
| Max output tokens | Not published | Not published | About 128K | 384K |
Availability
| Feature | GLM-5.3 | Kimi K3 | Qwen3.8 Max | DeepSeek V4 |
|---|---|---|---|---|
| Open weights | ✓ | Weight release announced, no date — not open yet | Claimed but unconfirmed | ✓ |
| Where you can run it | Z.ai API, GLM Coding Plan, ClinePass, OpenCode Go | Moonshot API, Kimi membership, ClinePass, OpenCode Go and Zen, Chutes | QwenCloud, DashScope, ClinePass, OpenCode Go, Token Plan | DeepSeek API, ClinePass, OpenCode Go |
| Released | 2026-08-14 | 2026-07-17 | About 2026-08-03 | V4-Flash-0731 and V4-Pro-0813 |
Verdict at a glance
| Feature | GLM-5.3 | Kimi K3 | Qwen3.8 Max | DeepSeek V4 |
|---|---|---|---|---|
| Best for | The cheapest credible agentic coding model with wide agent support | Long-context agentic work when you can get access | 1M context at half the price of the Western flagships | The lowest cost per token, and the largest output ceiling here |
The Verdict
Choose GLM-5.3 if...
- ✓You want the cheapest credible agentic model with the widest agent support
- ✓You want GLM-5.3-Flash at $0.15/$0.50 for the cheap half of your work
- ✓You are already on the GLM Coding Plan, ClinePass or OpenCode Go
- ✓Free cache storage during the current promotion matters at your volume
Choose Kimi K3 if...
- ✓Kimi K3's quality on your workload justifies $3/$15
- ✓You want a documented 1M context from a model available in five places
- ✓You want it in confidential compute — Chutes serves K3 at the full 1M
- ✓You accept that the weights are announced but not released
Choose Qwen3.8 Max if...
- ✓You want 1M context at $2/$6, half what the Western flagships charge
- ✓You want both OpenAI- and Anthropic-compatible endpoints from the vendor
- ✓Night-time credit burn at half rate on the Token Plan fits your schedule
- ✓About 128K output is enough for your longest generations
Choose DeepSeek V4 if...
- ✓You want the lowest cost per token of any credible flagship
- ✓You need the largest output ceiling here — 384K on a 1M context
- ✓You can schedule work outside the peak windows and pay half
- ✓Cache hits at $0.007-$0.044 make your repeated-context workload nearly free
- ✓You want genuinely open weights and a raw key with no tool restrictions
Frequently Asked Questions
What is the difference between GLM-5.3, Kimi K3, Qwen3.8 Max and DeepSeek V4?
GLM-5.3 is Z.ai's flagship, live on the API at $1.40 in and $4.40 out per million tokens — the same price as GLM-5.2 — with cache reads at $0.26 and cache storage free for a limited time. Kimi K3 is Moonshot's flagship at $3 in and $15 out per million tokens, cache reads at $0.30, with a 1M-token context. Qwen3.8 Max was released around 2026-08-03 at $2 in and $6 out per million tokens, with cache reads at $0.25 and explicit cache reads at $0.17. DeepSeek V4 is the most aggressively priced flagship here, and the only one that charges differently by time of day.
How much do GLM-5.3, Kimi K3, Qwen3.8 Max and DeepSeek V4 cost?
GLM-5.3: $1.40 in / $4.40 out per Mtok. Kimi K3: $3 in / $15 out per Mtok. Qwen3.8 Max: $2 in / $6 out per Mtok. DeepSeek V4: Flash $0.22/$0.66 off-peak, $0.44/$1.32 peak; Pro $0.66/$1.98 off-peak, $1.32/$3.96 peak. See the pricing table above for full plan details.
Which of GLM-5.3, Kimi K3, Qwen3.8 Max and DeepSeek V4 should I pick?
Choose GLM-5.3 if you want the cheapest credible agentic model with the widest agent support. Choose Kimi K3 if kimi K3's quality on your workload justifies $3/$15. Choose Qwen3.8 Max if you want 1M context at $2/$6, half what the Western flagships charge. Choose DeepSeek V4 if you want the lowest cost per token of any credible flagship.
Switch models without switching tools
Picking a model is a decision you will make again in three months, so the thing worth optimising is how cheaply you can change your mind. 1DevTool runs Claude Code, Codex CLI, Cline, OpenCode and Gemini CLI side by side in persistent terminals — point each at whichever model is winning this quarter, keep your own keys, and get an HTTP client, a 13-engine database client and an embedded browser in the same window. One-time $29, no subscription.