> 2026-07-08/09 · Cognition · AI coding model > Original post: https://cognition.com/blog/swe-1-7 > Secondary coverage: https://www.marktechpost.com/2026/07/09/cognition-ai-releases-swe-1-7
TL;DR
On July 8, Cognition (maker of Devin) released SWE-1.7 — an agentic software engineering model trained with large-scale asynchronous RL on top of the Kimi K2.7 base, an open-source model from Moonshot AI. The base model itself being a Chinese open-source LLM, and the engineering lessons that follow, are both worth unpacking.
Key numbers: frontier-level performance at far lower cost
The most striking part of the release is the cost-performance scatter plot: FrontierCode 1.1 Main scores (y-axis) vs. dollar cost per rollout (x-axis). SWE-1.7 sits on the Pareto frontier — near-frontier intelligence at a fraction of the cost of other frontier models.
FrontierCode 1.1 Main (100 high-quality agentic coding tasks)
- SWE-1.7: 42.3%
- Opus 4.8: 46.5% (+4.2 pp lead)
- GPT-5.5: 43.0% (+0.7 pp lead)
- Kimi K2.7 Code (base): 30.1% (12.2 pp below)
- SWE-1.6 (previous gen): 9.4% (over 4x improvement)
- SWE-1.7: 81.5% (Opus 4.8: 86.9%; GPT-5.5: 84.2%; Kimi K2.7 Code: 72.7%)
- SWE-1.7: 77.8% — above GPT-5.5's 76.8% (Opus 4.8: 84.4%)
- Network-isolated sandboxes (no web lookups for answers)
- Removed git history and reference artifacts
- Isolated grading path from the agent
- Programmatic checks for known exploit signatures
- Any cheating attempt is rewarded 0
- OpenAI GPT-5.6: proprietary closed base + massive capex + user choice complexity
- Cognition SWE-1.7: open-source base + application-layer RL + inference hardware innovation
Terminal-Bench 2.1
SWE-Bench Multilingual
Notably, SWE-1.7 beats GPT-5.5 on multilingual generalization despite its Kimi K2.7 base.
Engineering deep dive: why this RL run worked
1. Entropy-stabilized training
The biggest risk in RL training is entropy collapse — the rollout distribution narrows until only a few tokens get sampled and the training signal dies. Cognition keeps probability distributions consistent between rollout and trainer: the rollout side uses top-p sampling and records which tokens were actually kept; the trainer renormalizes probabilities using that record, avoiding KL divergence blowups.
A key detail: sampling distribution replay — replaying from the exact kept-token sets recorded at rollout, preventing drift like "this token had 30% probability at rollout but 5% at trainer time."
2. Cross-continent multi-cluster training
Training spanned 4 datacenters on 3 continents, with one trainer cluster in the US and rollout clusters worldwide, mixing in-house GPUs and third-party inference from Fireworks and others.
Syncing 1T-parameter weights across continents is hard engineering. Cognition's trick: transfer weight deltas, not full weights. Every K steps they compute the weight difference and relay it via cloud object storage (not point-to-point broadcast), cutting transfer volume by 99%+. The inference engine prefetches deltas into CPU memory and pauses only 3–4 seconds to apply them.
The result: cross-continental weight updates for a 1T-parameter model complete in 1–2 minutes — the infrastructure that makes asynchronous RL "continent-scale" possible.
3. Self-compaction for long tasks
Training rollouts ran up to 6 hours. The core challenge for such long-horizon tasks is context exhaustion. Cognition trained the model to generate its own summaries and resume from them, jointly optimizing (1) writing more concise summaries and (2) restoring work state from a summary.
They used an alternating length penalty: an Unconstrained phase optimizing only task success, and a Budget phase penalizing solutions exceeding token/turn/tool-call budgets. This compresses solved-task lengths while preserving long-horizon reasoning on hard tasks.
4. Data quality and anti-cheating
This is why FrontierCode 1.1 uses blocking criteria — failing solutions score 0 rather than partial credit.
Why this matters
1. The "post-training ceiling" hypothesis is refuted
Kimi K2.7 was already heavily RL-post-trained by Moonshot AI before Cognition got it. On top of that well-trained base, Cognition extracted +12.2 pp on FrontierCode and +8.8 pp on Terminal-Bench. This pushes back against the recent hypothesis that "frontier models are near the RL ceiling." Cognition's answer: with data, algorithms, and infrastructure all in place, RL still has significant headroom on already-trained bases. Good news for the agentic AI industry at large.
2. The "open-source base + top-tier RL fine-tune" model is validated
The business structure: open-source base from China (Kimi K2.7), RL fine-tuning and product by a US company (Cognition), delivered as SaaS (Devin). If it works, everyone wins: the base-model vendor gains adoption channel value, the third-party model company skips pretraining, and end customers get frontier capability at lower cost. Expect more "open base + application-layer RL" combos in the next 12 months.
3. Cerebras at 1000 TPS is a key enabler
SWE-1.7 launched day-one on Devin (Web/Desktop/CLI) via Cerebras at 1000 tokens per second. At 50 TPS, a 1M-token generation takes ~5.5 hours; at 1000 TPS, 17 minutes. Cerebras' wafer-scale engine is one of the few platforms that can run 1T-parameter models at this speed — this isn't "slightly cheaper" inference, it's inference fast enough to change the product's form.
4. Contrast with the same-week GPT-5.6 Sol launch
OpenAI released the GPT-5.6 family on July 9, one day after SWE-1.7 — but the routes differ:
Risks and open questions
1. FrontierCode 1.1 is Cognition's own benchmark. The 100 questions are public (plus 150 Extended), but ownership sits with Cognition. On the independent Terminal-Bench 2.1, SWE-1.7 trails GPT-5.5 by 2.7 pp. 2. Cerebras 1000 TPS is a commercial claim, not free — and may only hold at specific batch sizes. Production latency awaits real user feedback. 3. Base-model dependency risk. If Moonshot AI changes the Kimi K2.7 license or drops maintenance, Cognition's RL pipeline must be rebuilt — an inherent fragility of the open-base path. 4. Training cost. 4 datacenters across continents + 6-hour rollouts + custom MoE optimizer is not replicable by a small team. SWE-1.7's "low cost" is relative to GPT-5.5's inference pricing, not to a startup's cost of reproduction.
Overall, SWE-1.7 is a landmark event for AI coding: frontier intelligence doesn't require frontier pricing — provided an engineering team deeply optimizes the RL pipeline, inference hardware, and training infrastructure simultaneously.