Head-to-head

SWE-2 vs Grok 4.6

On the only table that scores both, SWE-2 leads Grok 4.6 on all four coding rows — by two to seven points. Then Terminal-Bench 4 puts them in the same tier, well behind the long-horizon frontier.

Updated September 12, 2026 · Independent comparison — SWE-2, Grok 4.6 are separate products.

S

SWE-2

Cognition's coding model, built for Devin

SWE-2 is Cognition's coding model, released September 10, 2026 and post-trained with reinforcement learning from Kimi K3 — Moonshot's 2.8T-parameter mixture-of-experts — in a single run that produces its medium, high and max effort levels together. It is co-designed with the Devin harness rather than sold as a general model: it shipped in Devin Desktop and Devin CLI, with Devin Web and Fusion following, and there is no public per-token API. Cognition's pitch is the cost curve. SWE-2 medium scores above SWE-1.7 while taking 58% fewer turns and costing 81% less, and reaches its first real edit in a median 18 steps against SWE-1.7's 48. The launch post puts it within one point of Fable 5.1 on FrontierCode 1.1 Main while being 64% cheaper.

G

Grok 4.6

xAI's frontier model, co-trained with Cursor

Grok 4.6 is the frontier model from xAI — branded SpaceXAI since the February 2026 acquisition — released August 12, 2026 and developed in collaboration with Cursor, which carries it as a first-party model and describes it as jointly trained. It is a general model built for coding, agentic work and knowledge tasks, with a 500K-token context window, a February 1, 2026 knowledge cutoff, and text and image input. Reasoning runs at low, medium, high (the default) and xhigh. On xAI's own launch table it scores 61 on the Artificial Analysis Intelligence Index, level with GPT-5.6 Sol Max and a point behind Fable 5 Max. Unlike SWE-2 you can simply buy it: the xAI API, Cursor, Grok Build's CLI, OpenRouter, Vercel and Cloudflare all carry it.

Bottom line

If you only read one number, read Terminal-Bench 4. On Cognition's September 10, 2026 launch table SWE-2 beats Grok 4.6 on all four coding rows — FrontierCode 1.1 Main 50.0% to 48.0%, DeepSWE 1.1 73.0% to 67.5%, Terminal-Bench 2.1 92.8% to 88.4%, Terminal-Bench 4 27.3% to 20.3% — so treating the two as roughly one class, with SWE-2 slightly ahead, is fair. But that same table has Fable 5.1 at 55.8% and GPT-6 Astra at 57.9% on Terminal-Bench 4, which means both of these models sit a long way outside the long-horizon frontier even though they look near-saturated on Terminal-Bench 2.1. The table is also vendor-run and mixes harnesses, and nobody has reproduced it. The durable difference is product shape: SWE-2 is a coding specialist with no public API and no published context window, reachable only through a Devin plan from $20/month. Grok 4.6 is a general model you can buy by the token at $2/$6, with a documented 500K window and first-party placement inside Cursor.

Cognition's launch table — the only grid scoring both

FeatureSWE-2Grok 4.6
FrontierCode 1.1 (Main split)50.0%48.0% (Cognition-reported)
DeepSWE 1.173.0%67.5% (Cognition-reported)
Terminal-Bench 2.192.8%88.4% (Cognition-reported)
Terminal-Bench 427.3%20.3% (Cognition-reported)
Same harness for both?No — SWE-2 ran in Devin CLINo — competitors ran in their own primary harnesses
Independently reproduced head-to-head

The Terminal-Bench 4 reality check

FeatureSWE-2Grok 4.6
Terminal-Bench 4 score27.3%20.3% (Cognition-reported)
Fable 5.1 on the same row55.8%55.8%
GPT-6 Astra on the same row57.9%57.9%
In the long-horizon frontier clusterNo — roughly 28 points backNo — roughly 36 points back
Near-saturated benches tell a different story92.8% on Terminal-Bench 2.188.4% on Terminal-Bench 2.1

What each lab reports on its own

FeatureSWE-2Grok 4.6
FrontierCode 1.1 ExtendedNot reported61.3% (xAI's own table)
DeepSWE — vendor's own figure73.0%65.9% high · 67.0% xhigh (xAI)
Terminal-Bench 3.0Not reported26.0% (xAI's own table)
CursorBench v3.2Not reported69.9% high · 70.8% xhigh
Artificial Analysis Intelligence IndexNot listed61 — ties GPT-5.6 Sol Max
SWE-bench Verified / ProNot publishedNot in xAI's announcement

Context, Price & Access

FeatureSWE-2Grok 4.6
Context windowNot published (base Kimi K3 is 1M — Cognition has not said SWE-2 inherits it)500K (xAI docs)
Knowledge cutoffNot publishedFebruary 1, 2026
Multimodal inputNot publishedText and image
Effort levelsMedium, High, Max — one RL runLow, Medium, High (default), xhigh
Per-token list priceNo public API rate; Cognition's Fusion post prints $0.75/Mtok for SWE-2 medium as an internal list figure$2 in / $6 out per 1M under 200K; $4 / $12 for the whole request at 200K and above
Cached inputNot published$0.50 per 1M (under 200K); $1.00 at or above
Cheapest way inDevin Pro at $20/mo — free in Desktop and CLI through October 10, 2026Pay-as-you-go on the xAI API
Standalone API key
Where you can run itDevin Desktop and Devin CLI; Devin Web and Fusion rolling outxAI API, Cursor (first-party), Grok Build CLI, Grok Bot, OpenRouter, Vercel, Cloudflare
Open weights
Base model disclosedKimi K3 (2.8T MoE, 104B active)Not disclosed — 1.5T-scale family per the model card

What each is built for

FeatureSWE-2Grok 4.6
Design targetAutonomous software engineering, co-designed with the Devin harnessGeneral frontier work — coding, agents and knowledge tasks
Use outside its vendor's productNo — Devin-locked, no public APIYes — API key, gateways and third-party IDEs
Pairs with another modelYes — Devin Fusion runs it as a sidekick under a Fable 5.1 or Astra leadNot a documented pattern
Efficiency claim58% fewer turns and 81% less cost than SWE-1.7, at a higher scoreFrontier-index parity with GPT-5.6 Sol Max at half the price of rival flagships

The Verdict

Choose SWE-2 if...

  • You are already paying for Devin — SWE-2 is the default there, and free in Desktop and CLI through October 10, 2026
  • Your work is agentic software engineering, not general reasoning, chat or vision
  • You want the better number on every row of the only table that scores both models
  • Cost per task matters more than cost per token — 58% fewer turns is the claim to test
  • You are willing to run inside one vendor's harness to get it

Choose Grok 4.6 if...

  • You want an API key and no product subscription wrapped around the model
  • You need a published, forecastable token price — $2/$6 under 200K, $4/$12 above it
  • You need one model for coding, chat and images rather than a coding specialist
  • A documented 500K context window matters more than an unpublished one
  • You work in Cursor, where Grok 4.6 is a first-party jointly-trained model rather than a metered third-party one
  • You want the widest set of places to run it — Cursor, Grok Build, OpenRouter, Vercel, Cloudflare

Frequently Asked Questions

Is SWE-2 as powerful as Grok 4.6?

On the only table that scores both, SWE-2 is slightly stronger. Cognition's launch post puts SWE-2 ahead of Grok 4.6 on all four of its coding rows — by 2.0 points on FrontierCode 1.1 Main, 5.5 on DeepSWE 1.1, 4.4 on Terminal-Bench 2.1 and 7.0 on Terminal-Bench 4. Treat it as a cluster rather than a blowout, and note the caveats: the table is Cognition's own, SWE-2 ran in Devin CLI while competitors ran in their own harnesses, and no independent party has reproduced it.

Does that mean SWE-2 is a frontier model?

Not on long-horizon agentic work. Cognition's own Terminal-Bench 4 row has SWE-2 at 27.3% and Grok 4.6 at 20.3%, against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra. Terminal-Bench 2.1 is near-saturated and shows a much flatter field at 92.8% versus 88.4%, which is why the choice of benchmark changes the story so much. Cognition's own framing is price-performance — near-frontier scores at a fraction of the cost — not outright frontier capability.

How much does SWE-2 cost per token?

There is no public per-token rate. SWE-2 is bundled into Devin's plans rather than sold as an API model, so the cheapest access is Devin Pro at $20/month — and Cognition is giving SWE-2 away free in Devin Desktop and the CLI through October 10, 2026. The one dollar figure Cognition has printed, $0.75 per million tokens for SWE-2 medium, appears in its Fusion post as an internal accounting price, not something you can buy. Grok 4.6 is straightforward by comparison: $2 in and $6 out per million tokens under a 200K prompt, doubling to $4 and $12 for the whole request at 200K and above.

What is SWE-2's context window?

Cognition has not published one. SWE-2 is post-trained from Kimi K3, whose base model documents a 1M-token window, but Cognition has never said SWE-2 inherits it and the launch post does not restate a window at all — its predecessor SWE-1.7 relied on self-compaction to run past the raw limit. Grok 4.6 publishes 500K directly in xAI's docs.

Can I use SWE-2 in Cursor, Windsurf or through an API?

No to all three. SWE-2 shipped in Devin Desktop and Devin CLI with Devin Web and Fusion following, and there is no standalone API key or gateway listing. Windsurf is not a separate option — that product was renamed Devin Desktop in June 2026, so SWE-2 in Windsurf is SWE-2 in Devin Desktop. Grok 4.6 is the opposite case: it is a first-party model inside Cursor, jointly trained with them, and also runs from the xAI API, Grok Build's CLI, OpenRouter, Vercel and Cloudflare.

Why do Grok 4.6's DeepSWE scores differ between sources?

Because the labs ran it themselves. Cognition's table reports Grok 4.6 at 67.5% on DeepSWE 1.1; xAI's own launch table reports 65.9% at high reasoning effort and 67.0% at xhigh. They are close but not identical, which is a useful reminder that vendor-run evals of a competitor's model carry harness and configuration choices the vendor made. The same caution applies to FrontierCode, where Cognition reports the Main split and xAI reports Extended — 50.0% and 61.3% are not comparable numbers.

Does either publish a SWE-bench Verified score?

Neither does, in the sources that matter. Cognition's SWE-2 launch reports FrontierCode, DeepSWE and Terminal-Bench instead — a change from the SWE-1.5 and 1.6 era, which did cite SWE-bench Pro. xAI's Grok 4.6 announcement and model docs do not include Verified either. Third-party leaderboards carry a 95.6% figure for Grok 4.6, but it is absent from xAI's own materials and SWE-bench Verified is widely considered saturated, so it is not used on this page.

1DevTool1DevTool

Run Devin and Grok side by side, on your own keys

SWE-2 only exists inside Devin and Grok 4.6 is strongest when you can point it anywhere — which is an argument for keeping both rather than picking one. 1DevTool runs Devin CLI, Grok Build, Claude Code, Codex, Cursor, Gemini CLI and twelve more agents side by side in persistent tmux-backed terminals, each on its own subscription or key, with automatic fallback when one hits a limit. Around them sits a 23-engine database client, an HTTP client and an embedded browser, all wired into Send-to-AI. From $29 one-time, no subscription and no credit pools.