Kimi K3 vs DeepSeek V4
Two Chinese-lab frontier models compared across size, multimodality, weights, pricing, and coding positioning.
Updated August 28, 2026 · Independent comparison — Kimi K3, DeepSeek V4 are separate products.
Kimi K3
Moonshot's 2.8T multimodal agentic model
Kimi K3 is Moonshot AI's 2.8-trillion-parameter model, released July 16, 2026 and billed as an "open 3T-class" system — though as of late August the weights have only been announced, with no release date, so it is API-first for now. It's a sparse Mixture-of-Experts (16 of 896 experts active) with native text, image, and video input, a 1M-token context, and a ~131K default output. Moonshot claims competitiveness with Anthropic's Fable 5, but publishes no independently verified benchmark table. It's served at $3.00/$15.00 per million tokens ($0.30 cache read) via Moonshot's API (model kimi-k3) and Kimi memberships, with a K3 Swarm Max variant for parallel multi-agent execution.
DeepSeek V4
Open-weight frontier family (Pro + Flash)
DeepSeek V4 is an open-weight family from April 24, 2026, split into V4-Pro (1.6T total / 49B active) and V4-Flash (284B / 13B). Both are Mixture-of-Experts with a 1M-token context (the default across DeepSeek's services) and up to 384K output, and both ship on Hugging Face and ModelScope. DeepSeek calls V4 open-source SOTA in agentic coding, claiming it beats all current open models in math, STEM, and coding, without publishing standalone SWE-bench numbers. Pricing is peak/off-peak: Pro runs $0.66/$1.98 off-peak and Flash just $0.22/$0.66, undercutting most closed frontier models, and it's reachable via OpenAI-, Anthropic-, and Responses-style endpoints.
Bottom line
Two Chinese-lab frontier models compared across size, multimodality, weights, pricing, and coding positioning. Choose Kimi K3 if you need native multimodal input (text, image, and video) in one model; choose DeepSeek V4 if you want a two-tier open family: a frontier Pro and a very cheap Flash ($0.22/$0.66 off-peak).
Model & Weights
| Feature | Kimi K3 | DeepSeek V4 |
|---|---|---|
| Developer | Moonshot AI | DeepSeek |
| Parameters | 2.8T total (sparse MoE, 16 of 896 experts) | Pro 1.6T/49B; Flash 284B/13B (MoE) |
| Open weights | Announced, no release date — API-first for now | Yes — HF + ModelScope (MIT per model card) |
| Multimodal input | Text, image, video | Text (experimental Flash-vision variant) |
| Variants | K3 + K3 Swarm Max | V4-Pro, V4-Flash |
Context & Output
| Feature | Kimi K3 | DeepSeek V4 |
|---|---|---|
| Context window | 1M tokens | 1M tokens (Pro & Flash) |
| Max output | ~131K default (configurable) | Up to 384K tokens |
| Reasoning / agent | Reasoning + agent workflows | Reasoning (V4) |
| Function calling | Yes (OpenAI-compatible) | ✓ |
Pricing (per 1M tokens)
| Feature | Kimi K3 | DeepSeek V4 |
|---|---|---|
| Input | $3.00 ($0.30 cache read) | $0.66-$1.32 (Pro) / $0.22-$0.44 (Flash) |
| Output | $15.00 | $1.98-$3.96 (Pro) / $0.66-$1.32 (Flash) |
| Pricing model | Flat + cache discount | Peak / off-peak by UTC window |
| Hosted access | Moonshot API, Kimi membership, ClinePass, OpenCode Go/Zen | DeepSeek API, ClinePass, OpenCode Go |
Coding & Benchmarks
| Feature | Kimi K3 | DeepSeek V4 |
|---|---|---|
| Published SWE-bench scores | Self-reported, not independently verified | Not published (numeric) |
| Vendor positioning | Competitive with Fable 5 (Moonshot claim) | Open-source SOTA in agentic coding (DeepSeek) |
| Native multimodality | ✓ | ✗ |
| Multi-agent execution | K3 Swarm Max | Not stated |
Access & Deployment
| Feature | Kimi K3 | DeepSeek V4 |
|---|---|---|
| Self-host weights | Not yet (weights unreleased) | Yes (~865GB Pro / ~160GB Flash) |
| Direct API | api.moonshot.ai (kimi-k3) | DeepSeek API + web (OpenAI/Anthropic/Responses) |
| Budget option | Single tier | V4-Flash (very low cost) |
| Released | July 16, 2026 | April 24, 2026 |
The Verdict
Choose Kimi K3 if...
- ✓You need native multimodal input (text, image, and video) in one model.
- ✓You want the largest model here (2.8T) and Moonshot's Fable-5-class positioning.
- ✓You run parallel multi-agent workflows (K3 Swarm Max).
- ✓You're fine with an API-first model and don't need downloadable weights today.
- ✓A steep cache-read discount ($0.30) fits your repeated-context workloads.
Choose DeepSeek V4 if...
- ✓You want a two-tier open family: a frontier Pro and a very cheap Flash ($0.22/$0.66 off-peak).
- ✓You want downloadable weights on Hugging Face and ModelScope to self-host or fine-tune.
- ✓You need up to 384K output tokens alongside a 1M-token context.
- ✓You can schedule heavy jobs into DeepSeek's off-peak window to cut costs.
- ✓You want lower hosted pricing than Kimi K3 on the flagship tier.
Frequently Asked Questions
What is the difference between Kimi K3 and DeepSeek V4?
Kimi K3 is Moonshot AI's 2.8-trillion-parameter model, released July 16, 2026 and billed as an "open 3T-class" system — though as of late August the weights have only been announced, with no release date, so it is API-first for now. DeepSeek V4 is an open-weight family from April 24, 2026, split into V4-Pro (1.6T total / 49B active) and V4-Flash (284B / 13B).
How much do Kimi K3 and DeepSeek V4 cost?
Kimi K3: $3.00 ($0.30 cache read). DeepSeek V4: $0.66-$1.32 (Pro) / $0.22-$0.44 (Flash). See the pricing table above for full plan details.
Should I choose Kimi K3 or DeepSeek V4?
Choose Kimi K3 if you need native multimodal input (text, image, and video) in one model. Choose DeepSeek V4 if you want a two-tier open family: a frontier Pro and a very cheap Flash ($0.22/$0.66 off-peak).
Wire either model into your agent
Kimi K3 and DeepSeek V4 are both OpenAI-compatible (DeepSeek is also open-weight and self-hostable). 1DevTool runs Claude Code, Codex CLI, Gemini CLI, Cline, and OpenCode side by side in persistent terminals — point them at a Moonshot key, a DeepSeek key, or your own vLLM server, switch per task, and use the built-in HTTP client, 13-engine database client, and embedded browser. One-time $29, no subscription.