A dedicated chip normally takes 18 to 36 months from architecture definition to tape-out. On July 18, 2026, Moonshot AI released Kimi K3 with a demo: a model-driven agent ran autonomously for 48 consecutive hours, using only open-source EDA tools and the Nangate 45nm process library to complete the design, optimization, and verification of an inference accelerator prototype. Just over a month later, on August 25 at Hot Chips, OpenAI published the first measured data for its in-house inference chip Jalapeño — TSMC 3nm, co-developed with Broadcom, nine months from architecture to tape-out. Both events point to the same thing: the capability boundary of AI coding is moving from "writing application software" toward "building physical hardware."
⏱️ Two Timelines: 48 Hours and 9 Months
Start with the two hardest numbers.
On the Kimi K3 line, the official framing is "early proof of concept." 48 hours of continuous autonomous operation, with an open-source EDA toolchain and the Nangate 45nm cell library as input, and an inference accelerator prototype capable of running Kimi's own Nano model as output.
The Jalapeño line is a real chip. Project started mid-2024, taped out in November 2025, and first measured data delivered on August 25, 2026 — under two years end to end. Engineering samples are already running OpenAI's own models in the lab, with small-scale deployment planned for late 2026 and scaling in 2027.
The two numbers aren't on the same scale. Kimi K3 used a 45nm open-source library, several generations behind the current 3nm/2nm frontier, and the company itself stated there was no tape-out, no silicon, no test board, and no independent reproduction. Jalapeño is real silicon on TSMC N3P, already back and running. Reading this as a "who's faster" contest misses the point — these are two different routes.
🔬 Kimi K3's Chip: 1.46 Million Standard Cells in 3.98 mm²
The late-July technical report laid out the details:
| Dimension | Value | |---|---| | Process library | Nangate 45nm open-source cell library | | Area | ~3.981 mm² (often reported as 4 mm²) | | Modules | 13 | | Standard cells | ~1.46 million | | On-chip memory | 0.277 MB SRAM | | Compute array | INT4 MAC array with on-chip dequantization | | Timing | Closed at 100 MHz | | Simulated throughput | 8,721 tokens/s |
The model itself is a 2.8-trillion-parameter MoE with a 1M-token context window, activating 16 of 896 experts at inference. Beyond the chip, it also wrote a Triton-like GPU compiler from scratch — MiniTriton (an MLIR-based tile-level IR, optimization passes, PTX code generation) — matching or beating official Triton and torch.compile on Roofline benchmarks.
💥 Aftermath of the 48 Hours: EDA Stocks Dropped for a Day
On the first trading day after the announcement, Synopsys fell over 12% intraday before closing down about 8%; Cadence fell 7.85% on the day. The panic spread along the supply chain: Zhipu fell 28% in Hong Kong, MiniMax fell 16%, and SoftBank fell 9% in Tokyo.
Sell-side interpretations split into two camps. Union Bancaire Privée managing director Vey-Sern Ling argued that if US companies shifted to Chinese models and bought fewer Anthropic services, capex would shrink and eventually hit chip demand. Morgan Stanley analyst Gary Yu framed it as the result of long-term cumulative progress in China's AI industry, not an overnight disruption. Bloomberg Intelligence analyst Robin Zhu considered the sell-off a reasonable reaction.
A month later, the narrative reversed. When OpenAI revealed Jalapeño's measured data, even as "LLMs topple EDA" talk peaked, Synopsys and Cadence stocks actually rose. A Yicain (First Financial) retrospective on September 9 explained why: K3 attempted to prove "AI can work without the EDA giants' tools," while Jalapeño — a chip with real deployment — remains deeply bound to commercial EDA, with low-level physical verification, sign-off, and advanced packaging still done by commercial and deterministic tools. The market switched from "replacement panic" to "assistive value-add."
🔥 Jalapeño's Measured Data: 1.5–1.9x per Watt, 1.7–3.6x Lower Latency
Tests used SemiAnalysis's public InferenceMAX benchmark, comparing against NVIDIA GB200 and GB300 systems across three open-weight models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T.
| Metric | Jalapeño result | |---|---| | Peak AI throughput per watt | 1.5–1.9x the comparison systems | | End-to-end latency | 1.7–3.6x lower (~28%–59% of the comparison) | | High-interactivity workloads | 2.1–4.1x higher | | Kimi K2.5 1T | ~1.5x better per watt, 3.4x lower latency | | Rated power | 700 W, sustained under 550 W in test loads | | Memory | 6 stacks HBM4, 216 GB total, 15.4 TB/s bandwidth |
The comparison GB300 has 288 GB HBM3E and a 1,400 W rated power. OpenAI hardware VP Richard Ho said AI systems usually trade off latency against throughput, and Jalapeño aims to have both.
Four caveats must be stated together. The comparison excludes Rubin; the data was mainly provided by OpenAI, with SemiAnalysis verifying some InferenceMAX runs on-site in the lab but not completing the full suite or the AgentX results closer to real agentic production loads; the tests didn't cover all production workloads; and Jalapeño used single-token prediction while some of NVIDIA's public production data used multi-token prediction — SemiAnalysis believes the margin could change under uniform settings.
🤖 The More Interesting Part: AI Writing Kernels for the Chip
More interesting than the performance charts is how Jalapeño came to be.
Engineers wrote kernels in OpenAI's in-house programming language Gluon; a kernel could run to thousands of lines and required repeated manual tuning against the chip's structure. OpenAI then put internal Codex together with GPT-Astra on the task, and within two months ported three open-weight models that weren't in the production plan onto Jalapeño. In some GPT-OSS attention and MoE modules, AI-generated implementations ran 1.5–1.8x faster than the original human-expert versions (OpenAI explicitly noted this figure applies only to selected modules and cannot be extrapolated to whole models). Other reports mention a 56% area optimization for BF16 multipliers.
This loop touches something deeper than performance. ASICs have always delivered strong performance for specific workloads, but every new model means rewriting kernels, adapting operators, and maintaining compilers and debug tooling. GPUs are expensive, and CUDA's value lies in spreading that engineering cost across a massive developer ecosystem. What OpenAI demonstrated is a different path: instead of replicating CUDA wholesale, let models automatically generate, test, and optimize low-level code, driving down the adaptation cost of getting new models onto in-house silicon.
This path is internal-only for now. Jalapeño primarily serves OpenAI's own models and workloads; external developers can't yet deploy and maintain applications on the platform the way they can with CUDA. It shows OpenAI can reduce CUDA dependence internally — not that OpenAI has built a market-facing alternative ecosystem.
🧭 The NVIDIA Relationship Isn't Over
Two weeks before the data drop, on August 17, NVIDIA announced credit support for SB Energy's PORTS-Pike campus in Ohio, reportedly up to $105 billion per the Financial Times and others. This isn't NVIDIA lending directly to OpenAI: SB Energy builds and operates, OpenAI signs a 20-year lease, and the campus will exclusively use NVIDIA's compute platform.
Jalapeño does inference only, not training. Richard Ho told Bloomberg: "NVIDIA is a great partner, and we still need a lot of NVIDIA." OpenAI and Broadcom announced a 10 GW multi-generation custom accelerator partnership in October 2025, targeting deployment starting in H2 2026 and completion by end-2029. The second-generation Jalapeño is in late-stage development and expected to tape out within months; third-generation concept design has begun.
⚖️ Three Hurdles Not Yet Cleared
| Unverified item | What it is | |---|---| | Kimi K3's chip | 45nm open-source library, no tape-out, no silicon, no test board, no independent reproduction; officially an early proof of concept | | Jalapeño's math | Performance-per-watt uses publicly stated package power from both sides, not same-lab measured draw; the full accounting of mass production, yield, HBM and packaging costs, and long-term operations isn't done | | Software ecosystem | Every new model, operator, and runtime mode requires further software-stack adaptation; Jalapeño hasn't yet proven this long-term software bill shrinks |
There's also a more basic boundary: the three test models are large but not the latest open models from their respective labs. How quickly newer, more complex models adapt to Jalapeño — and how they perform — remains to be seen.
🔁 Where This Sits in AI Coding History
The past two years of AI coding centered on "can a model handle a real repo-level change," with benchmarks moving from single-file completion to SWE-bench to long-horizon autonomous tasks. Kimi K3's 48 hours and OpenAI's nine months push the line one notch further: when an agent can work continuously for two days without collapsing and iterate through verification loops, what it can take on isn't just software repos, but physical design flows described by deterministic toolchains.
But there's a hard constraint on this reasoning. Chip design involves extensive physical verification, sign-off, and advanced packaging that rely on commercial EDA's deterministic tools — places where a language model can't substitute on probability. Jalapeño's nine-month tape-out owed just as much to Broadcom's deep involvement and the commercial EDA foundation.
Three numbers to watch over the next 12 months: the real inference-cost reduction after Jalapeño deploys at scale, whether gen-2 tape-out timing compresses further, and the labor-substitution ratio of Codex-like tools in chip kernel adaptation.
📚 References
1. Yicain/Sina Finance: LLMs enter semiconductor design — will EDA "tremble"? 2026-09-09 2. NetEase Smart/Tencent News: OpenAI's first chip stuns, gen one reaches the frontier, 2026-08-26 3. Zhi Dongxi: Kimi K3 technical report and weight release, 2026-07-28 4. Anue: Tape-out in 9 months — OpenAI's custom ASIC arrives, 2026-08 5. SemiAnalysis InferenceMAX benchmark / Hot Chips 2026 (2026-08-25) 6. MartechAI: OpenAI's Jalapeño AI Chip Shows Higher Speed and Power Efficiency in First Benchmarks