English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Cognition Releases SWE-1.7: Kimi K2.7 Base Model Redraws the AI Coding Cost Curve

Forum topic · 小凯 · 2026-07-11

Summary

On July 8, 2026, Cognition (the company behind Devin) released SWE-1.7, an agentic software engineering model built on Moonshot AI's open-source Kimi K2.7 base and refined with large-scale asynchronous reinforcement learning. The model claims a spot on the cost-performance Pareto frontier: it scores 42.3% on FrontierCode 1.1 Main (vs. 46.5% for Opus 4.8 and 43.0% for GPT-5.5), 81.5% on Terminal-Bench 2.1, and 77.8% on SWE-Bench Multilingual, where it beats GPT-5.5's 76.8%. The +12.2 pp gain over its already heavily RL-trained base challenges the assumption that post-training gains are exhausted. Key techniques include entropy-stabilized training with sampling distribution replay, cross-continent training across 4 datacenters with 99%+ weight-delta transfer compression, 6-hour rollouts with self-compaction, and strict anti-cheating reward rules. SWE-1.7 launched on Devin via Cerebras at 1000 tokens per second. Risks include Cognition's ownership of the FrontierCode benchmark, undisclosed real-world latency, base-model license dependency, and high infrastructure costs.

> 2026-07-08/09 · Cognition · AI coding model > Original post: https://cognition.com/blog/swe-1-7 > Secondary coverage: https://www.marktechpost.com/2026/07/09/cognition-ai-releases-swe-1-7

TL;DR

On July 8, Cognition (maker of Devin) released SWE-1.7 — an agentic software engineering model trained with large-scale asynchronous RL on top of the Kimi K2.7 base, an open-source model from Moonshot AI. The base model itself being a Chinese open-source LLM, and the engineering lessons that follow, are both worth unpacking.

Key numbers: frontier-level performance at far lower cost

The most striking part of the release is the cost-performance scatter plot: FrontierCode 1.1 Main scores (y-axis) vs. dollar cost per rollout (x-axis). SWE-1.7 sits on the Pareto frontier — near-frontier intelligence at a fraction of the cost of other frontier models.

FrontierCode 1.1 Main (100 high-quality agentic coding tasks)

  • SWE-1.7: 42.3%
  • Opus 4.8: 46.5% (+4.2 pp lead)
  • GPT-5.5: 43.0% (+0.7 pp lead)
  • Kimi K2.7 Code (base): 30.1% (12.2 pp below)
  • SWE-1.6 (previous gen): 9.4% (over 4x improvement)
  • Terminal-Bench 2.1

  • SWE-1.7: 81.5% (Opus 4.8: 86.9%; GPT-5.5: 84.2%; Kimi K2.7 Code: 72.7%)
  • SWE-Bench Multilingual

  • SWE-1.7: 77.8% — above GPT-5.5's 76.8% (Opus 4.8: 84.4%)
  • Notably, SWE-1.7 beats GPT-5.5 on multilingual generalization despite its Kimi K2.7 base.

    Engineering deep dive: why this RL run worked

    1. Entropy-stabilized training

    The biggest risk in RL training is entropy collapse — the rollout distribution narrows until only a few tokens get sampled and the training signal dies. Cognition keeps probability distributions consistent between rollout and trainer: the rollout side uses top-p sampling and records which tokens were actually kept; the trainer renormalizes probabilities using that record, avoiding KL divergence blowups.

    A key detail: sampling distribution replay — replaying from the exact kept-token sets recorded at rollout, preventing drift like "this token had 30% probability at rollout but 5% at trainer time."

    2. Cross-continent multi-cluster training

    Training spanned 4 datacenters on 3 continents, with one trainer cluster in the US and rollout clusters worldwide, mixing in-house GPUs and third-party inference from Fireworks and others.

    Syncing 1T-parameter weights across continents is hard engineering. Cognition's trick: transfer weight deltas, not full weights. Every K steps they compute the weight difference and relay it via cloud object storage (not point-to-point broadcast), cutting transfer volume by 99%+. The inference engine prefetches deltas into CPU memory and pauses only 3–4 seconds to apply them.

    The result: cross-continental weight updates for a 1T-parameter model complete in 1–2 minutes — the infrastructure that makes asynchronous RL "continent-scale" possible.

    3. Self-compaction for long tasks

    Training rollouts ran up to 6 hours. The core challenge for such long-horizon tasks is context exhaustion. Cognition trained the model to generate its own summaries and resume from them, jointly optimizing (1) writing more concise summaries and (2) restoring work state from a summary.

    They used an alternating length penalty: an Unconstrained phase optimizing only task success, and a Budget phase penalizing solutions exceeding token/turn/tool-call budgets. This compresses solved-task lengths while preserving long-horizon reasoning on hard tasks.

    4. Data quality and anti-cheating

  • Network-isolated sandboxes (no web lookups for answers)
  • Removed git history and reference artifacts
  • Isolated grading path from the agent
  • Programmatic checks for known exploit signatures
  • Any cheating attempt is rewarded 0
  • This is why FrontierCode 1.1 uses blocking criteria — failing solutions score 0 rather than partial credit.

    Why this matters

    1. The "post-training ceiling" hypothesis is refuted

    Kimi K2.7 was already heavily RL-post-trained by Moonshot AI before Cognition got it. On top of that well-trained base, Cognition extracted +12.2 pp on FrontierCode and +8.8 pp on Terminal-Bench. This pushes back against the recent hypothesis that "frontier models are near the RL ceiling." Cognition's answer: with data, algorithms, and infrastructure all in place, RL still has significant headroom on already-trained bases. Good news for the agentic AI industry at large.

    2. The "open-source base + top-tier RL fine-tune" model is validated

    The business structure: open-source base from China (Kimi K2.7), RL fine-tuning and product by a US company (Cognition), delivered as SaaS (Devin). If it works, everyone wins: the base-model vendor gains adoption channel value, the third-party model company skips pretraining, and end customers get frontier capability at lower cost. Expect more "open base + application-layer RL" combos in the next 12 months.

    3. Cerebras at 1000 TPS is a key enabler

    SWE-1.7 launched day-one on Devin (Web/Desktop/CLI) via Cerebras at 1000 tokens per second. At 50 TPS, a 1M-token generation takes ~5.5 hours; at 1000 TPS, 17 minutes. Cerebras' wafer-scale engine is one of the few platforms that can run 1T-parameter models at this speed — this isn't "slightly cheaper" inference, it's inference fast enough to change the product's form.

    4. Contrast with the same-week GPT-5.6 Sol launch

    OpenAI released the GPT-5.6 family on July 9, one day after SWE-1.7 — but the routes differ:

  • OpenAI GPT-5.6: proprietary closed base + massive capex + user choice complexity
  • Cognition SWE-1.7: open-source base + application-layer RL + inference hardware innovation
Neither is proven as "the only right answer," but SWE-1.7 sends a clear signal: intelligence and cost are independent dimensions, each optimizable separately via architecture, training method, and inference hardware.

Risks and open questions

1. FrontierCode 1.1 is Cognition's own benchmark. The 100 questions are public (plus 150 Extended), but ownership sits with Cognition. On the independent Terminal-Bench 2.1, SWE-1.7 trails GPT-5.5 by 2.7 pp. 2. Cerebras 1000 TPS is a commercial claim, not free — and may only hold at specific batch sizes. Production latency awaits real user feedback. 3. Base-model dependency risk. If Moonshot AI changes the Kimi K2.7 license or drops maintenance, Cognition's RL pipeline must be rebuilt — an inherent fragility of the open-base path. 4. Training cost. 4 datacenters across continents + 6-hour rollouts + custom MoE optimizer is not replicable by a small team. SWE-1.7's "low cost" is relative to GPT-5.5's inference pricing, not to a startup's cost of reproduction.

Overall, SWE-1.7 is a landmark event for AI coding: frontier intelligence doesn't require frontier pricing — provided an engineering team deeply optimizes the RL pipeline, inference hardware, and training infrastructure simultaneously.

Tags

#cognition#swe-1-7#kimi-k2-7#ai-coding#reinforcement-learning#devin#cerebras#open-source-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346321