4-way comparison

GLM-5.3 vs Kimi K3 vs Qwen3.8 Max vs DeepSeek V4

The four flagships behind almost every cheap coding pass. Two have genuinely open weights, one has an announcement, and one is unconfirmed.

Updated August 28, 2026 · Independent comparison — GLM-5.3, Kimi K3, Qwen3.8 Max, DeepSeek V4 are separate products.

G

GLM-5.3

Z.ai's flagship, launched 2026-08-14

GLM-5.3 is Z.ai's flagship, live on the API at $1.40 in and $4.40 out per million tokens — the same price as GLM-5.2 — with cache reads at $0.26 and cache storage free for a limited time. Z.ai does not publish an official context figure for 5.3. It reaches developers through the Z.ai API, the GLM Coding Plan (which auto-routes 5.2 and 5.1 requests up to 5.3), ClinePass and OpenCode Go, where the cheaper GLM-5.3-Flash runs at $0.15/$0.50.

K

Kimi K3

Moonshot's 1M-context flagship

Kimi K3 is Moonshot's flagship at $3 in and $15 out per million tokens, cache reads at $0.30, with a 1M-token context. It is reachable through the Moonshot API, the Kimi membership from the Moderato tier (though the full 1M chat context is gated to the top tier), ClinePass, OpenCode Go and Zen, and Chutes in confidential compute at the full 1M. A weight release has been announced but has no date, so K3 is not yet an open-weight model.

Q

Qwen3.8 Max

Alibaba's flagship, released August 2026

Qwen3.8 Max was released around 2026-08-03 at $2 in and $6 out per million tokens, with cache reads at $0.25 and explicit cache reads at $0.17. Context is 1M with roughly 128K output. It is served from QwenCloud and DashScope on both OpenAI- and Anthropic-compatible endpoints, and appears in ClinePass, OpenCode Go and the QwenCloud Token Plan, where night-time usage between 22:00 and 08:00 UTC+8 burns credits at half rate. Open weights were claimed shortly after launch but remain unconfirmed.

D

DeepSeek V4

1M context, 384K output, peak/off-peak pricing

DeepSeek V4 is the most aggressively priced flagship here, and the only one that charges differently by time of day. Flash costs $0.22 in and $0.66 out per million tokens off-peak, doubling during peak hours (Mon-Fri 01:00-04:00 and 06:00-10:00 UTC); Pro runs $0.66/$1.98 off-peak and $1.32/$3.96 at peak. Cache hits are close to free at $0.007-$0.044. It offers a 1M context with up to 384K output — the largest output ceiling in this group — across OpenAI-, Anthropic- and Responses-compatible endpoints.

Bottom line

These four models sit behind almost every cheap coding pass, and they are not equally open. GLM-5.3 is the cheapest credible option at $1.40/$4.40 with the widest agent support, though Z.ai does not publish its context window. DeepSeek V4 is cheaper still and uniquely time-of-day priced — Flash from $0.22/$0.66 off-peak — with the largest output ceiling here at 384K on a 1M context, and near-free cache hits. Qwen3.8 Max is $2/$6 with a 1M context and about 128K output; open weights were claimed shortly after its August launch but remain unconfirmed. Kimi K3 is the most expensive at $3/$15, offers 1M context, and despite an announced weight release has no date — so it is not an open-weight model today.

Pricing

FeatureGLM-5.3Kimi K3Qwen3.8 MaxDeepSeek V4
Price per Mtok (in / out)$1.40 in / $4.40 out per Mtok$3 in / $15 out per Mtok$2 in / $6 out per MtokFlash $0.22/$0.66 off-peak, $0.44/$1.32 peak; Pro $0.66/$1.98 off-peak, $1.32/$3.96 peak
Cache pricing$0.26 read; storage free for a limited time$0.30 read$0.25 read; $0.17 explicit cache read$0.007-$0.044 per Mtok — effectively free

Capacity

FeatureGLM-5.3Kimi K3Qwen3.8 MaxDeepSeek V4
Context windowNot published1M1M1M
Max output tokensNot publishedNot publishedAbout 128K384K

Availability

FeatureGLM-5.3Kimi K3Qwen3.8 MaxDeepSeek V4
Open weightsWeight release announced, no date — not open yetClaimed but unconfirmed
Where you can run itZ.ai API, GLM Coding Plan, ClinePass, OpenCode GoMoonshot API, Kimi membership, ClinePass, OpenCode Go and Zen, ChutesQwenCloud, DashScope, ClinePass, OpenCode Go, Token PlanDeepSeek API, ClinePass, OpenCode Go
Released2026-08-142026-07-17About 2026-08-03V4-Flash-0731 and V4-Pro-0813

Verdict at a glance

FeatureGLM-5.3Kimi K3Qwen3.8 MaxDeepSeek V4
Best forThe cheapest credible agentic coding model with wide agent supportLong-context agentic work when you can get access1M context at half the price of the Western flagshipsThe lowest cost per token, and the largest output ceiling here

The Verdict

Choose GLM-5.3 if...

  • You want the cheapest credible agentic model with the widest agent support
  • You want GLM-5.3-Flash at $0.15/$0.50 for the cheap half of your work
  • You are already on the GLM Coding Plan, ClinePass or OpenCode Go
  • Free cache storage during the current promotion matters at your volume

Choose Kimi K3 if...

  • Kimi K3's quality on your workload justifies $3/$15
  • You want a documented 1M context from a model available in five places
  • You want it in confidential compute — Chutes serves K3 at the full 1M
  • You accept that the weights are announced but not released

Choose Qwen3.8 Max if...

  • You want 1M context at $2/$6, half what the Western flagships charge
  • You want both OpenAI- and Anthropic-compatible endpoints from the vendor
  • Night-time credit burn at half rate on the Token Plan fits your schedule
  • About 128K output is enough for your longest generations

Choose DeepSeek V4 if...

  • You want the lowest cost per token of any credible flagship
  • You need the largest output ceiling here — 384K on a 1M context
  • You can schedule work outside the peak windows and pay half
  • Cache hits at $0.007-$0.044 make your repeated-context workload nearly free
  • You want genuinely open weights and a raw key with no tool restrictions

Frequently Asked Questions

What is the difference between GLM-5.3, Kimi K3, Qwen3.8 Max and DeepSeek V4?

GLM-5.3 is Z.ai's flagship, live on the API at $1.40 in and $4.40 out per million tokens — the same price as GLM-5.2 — with cache reads at $0.26 and cache storage free for a limited time. Kimi K3 is Moonshot's flagship at $3 in and $15 out per million tokens, cache reads at $0.30, with a 1M-token context. Qwen3.8 Max was released around 2026-08-03 at $2 in and $6 out per million tokens, with cache reads at $0.25 and explicit cache reads at $0.17. DeepSeek V4 is the most aggressively priced flagship here, and the only one that charges differently by time of day.

How much do GLM-5.3, Kimi K3, Qwen3.8 Max and DeepSeek V4 cost?

GLM-5.3: $1.40 in / $4.40 out per Mtok. Kimi K3: $3 in / $15 out per Mtok. Qwen3.8 Max: $2 in / $6 out per Mtok. DeepSeek V4: Flash $0.22/$0.66 off-peak, $0.44/$1.32 peak; Pro $0.66/$1.98 off-peak, $1.32/$3.96 peak. See the pricing table above for full plan details.

Which of GLM-5.3, Kimi K3, Qwen3.8 Max and DeepSeek V4 should I pick?

Choose GLM-5.3 if you want the cheapest credible agentic model with the widest agent support. Choose Kimi K3 if kimi K3's quality on your workload justifies $3/$15. Choose Qwen3.8 Max if you want 1M context at $2/$6, half what the Western flagships charge. Choose DeepSeek V4 if you want the lowest cost per token of any credible flagship.

1DevTool1DevTool

Switch models without switching tools

Picking a model is a decision you will make again in three months, so the thing worth optimising is how cheaply you can change your mind. 1DevTool runs Claude Code, Codex CLI, Cline, OpenCode and Gemini CLI side by side in persistent terminals — point each at whichever model is winning this quarter, keep your own keys, and get an HTTP client, a 13-engine database client and an embedded browser in the same window. One-time $29, no subscription.