English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Jim Keller's Tenstorrent Bet: Can Open-Source RISC-V AI Chips Challenge NVIDIA?

Forum topic · 小凯 · 2026-06-07

Summary

This article analyzes Jim Keller's Tenstorrent and its bid to challenge NVIDIA's AI inference dominance using open RISC-V-based silicon and an open-source software stack. It recounts Keller's track record (AMD K7/K8 and Zen, Apple A4/A5, Tesla FSD, Intel Xe) and details Tenstorrent's Tensix cores: five RISC-V processors per core, no hardware cache hierarchy, software-controlled data movement, and on-chip Ethernet mesh interconnect. The company's differentiation lies in its fully open stack (TT-Forge compiler, TT-Metalium SDK, MIT licensing) and a roadmap from Grayskull to Blackhole (480 Tensix cores) and the chiplet-based Grendel. Official claims like 350+ tokens/second on DeepSeek R1 and 5x lower TCO versus GB300 are examined skeptically given undisclosed test conditions and software immaturity. The piece also covers NVIDIA's reported ~$20B Groq acquisition as a cautionary precedent, Tenstorrent's three-pronged business model (hardware sales, IP licensing to LG/Hyundai/Japan's LSTC, sovereign AI deals), its $2.6B Series D valuation, and the broader wave of NVIDIA challengers including Cerebras, Etched, and hyperscaler in-house chips. Conclusion: no short-term displacement of NVIDIA, but long-term potential in a fragmented, open-architecture market.

Jim Keller's Tenstorrent Bet: Can Open-Source RISC-V Chips Topple NVIDIA's Empire?

> A note on the source: a video title misspells the company as "TenStorent" — itself a sign that Tenstorrent's chips aren't yet big enough on the table for everyone to know its name. But Jim Keller's résumé commands attention: AMD Zen, Apple A4/A5, Tesla FSD, Intel Xe. The 64-year-old chip legend is now making the last big bet of his career: using open-source RISC-V architecture to take AI inference market share from NVIDIA. The stakes are real — in December 2025, NVIDIA reportedly spent ~$20 billion on Groq's LPU technology and team, a move the author frames as an "execution" rather than an acquisition. Does Tenstorrent have something NVIDIA can't buy?

Key points

  • Jim Keller's track record: AMD K7/K8 (first real threat to Intel in x86), Apple A4/A5 (foundation of the iPhone/iPad), AMD Zen (rescued AMD), Tesla FSD (first automotive AI inference chip), Intel Xe (unsuccessful), now Tenstorrent.
  • Philosophy: don't optimize what others have optimized — redefine the problem. Tenstorrent is his hardest hand yet against a $3-trillion, full-stack NVIDIA.
  • Tenstorrent's technical approach: neither GPU nor TPU

    Tensix cores

  • Each Tensix core contains 5 RISC-V processors plus vector/matrix units and local SRAM.
  • No hardware cache hierarchy (L1/L2/L3); data movement is explicitly software-controlled.
  • Chips connect via on-chip Ethernet in a mesh/torus topology — no NVLink or external switches required.
  • | Dimension | NVIDIA GPU | Tenstorrent Tensix | |---|---|---| | Control | SIMT | 5 independent RISC-V cores | | Memory | Hardware-managed caches | Explicit SRAM management, no caches | | Interconnect | NVLink/InfiniBand + switches | On-chip Ethernet, direct torus | | Software | CUDA (proprietary) | TT-Forge/TT-Metalium (open source) | | Scaling | Dedicated interconnect hardware | Standard Ethernet |

    Keller's bet: AI dataflow patterns are predictable, so transistors spent on caches, branch prediction, and out-of-order execution are wasted — better to spend them on matrix compute and let software move data precisely.

    Fully open-source software stack

  • TT-Forge: open AI compiler supporting PyTorch/JAX/ONNX (public beta)
  • TT-Metalium: low-level SDK for writing kernels, MIT license
  • TT-LLK: low-level kernel software
  • RISC-V ISA: royalty-free open instruction set
  • The strategy: replicate Linux's victory over Unix — use open-ecosystem long-term stickiness to counter CUDA's proprietary short-term advantages.

    Chip roadmap

    | Generation | Timing | Notes | |---|---|---| | Grayskull | 2020 | Proof of concept | | Wormhole | 2022–2023 | Better interconnect, training + inference | | Blackhole | 2024–2026 | Current flagship: 480 Tensix cores, 2,654 TFLOPS (BlockFP8) | | Grendel | 2026+ | Chiplet architecture (Open Chiplet Atlas), separate CPU and AI tiles |

    Performance claims vs. reality

    Tenstorrent's May 2026 TT-Deploy event claimed:

  • Galaxy Blackhole server: DeepSeek R1 inference at 350+ tokens/second
  • TCO 5x lower than NVIDIA GB300
  • "Blitz Mode": 10x faster generative AI video vs. current GPUs
  • Keller: "We are committed to crushing everybody at everything"
  • The article urges skepticism: undisclosed test conditions (batch size, quantization, prefill/decode split) make token-rate claims hard to verify; TCO figures may ignore software migration and retraining costs; the "Blitz Mode" comparison baseline is unclear; and TT-Forge's beta status means production readiness remains unproven.

    Galaxy Blackhole vs. GB300

    | Dimension | NVIDIA GB300 | Galaxy Blackhole | |---|---|---| | Memory | HBM3e | GDDR6 / on-chip SRAM | | Ecosystem | CUDA (20 years) | TT-Forge (beta) | | Cloud availability | All major clouds | Koyeb and partners | | Model support | Nearly all AI models | Compile-dependent, limited | | Production validation | Massive global deployment | Small deployments, Japan/Korea sovereign projects |

    Bottom line: Tenstorrent may win on specific workloads, but generality and ecosystem maturity trail NVIDIA by orders of magnitude — a 5-year startup vs. a 20-year incumbent.

    The $20B Groq precedent

  • Groq's LPU (designed by Google TPU father Jonathan Ross) removes caches, branch prediction, and OoO execution entirely, using compiler-deterministic execution and on-chip SRAM — hitting 241 tokens/s on Llama 2 70B in early 2024.
  • On December 24, 2025, NVIDIA announced a ~$20 billion deal for Groq's core technology and team (including Ross and Sunny Madra) — nearly 3x the Mellanox deal.
  • Motives: GPUs are weak at the memory-bandwidth-bound decode phase where Groq's SRAM architecture excels; buying Groq removes a threat and absorbs key talent.
  • Lesson for Tenstorrent: novel inference architectures have value, but the ceiling may be an acquisition price. Keller's counter is open source — NVIDIA can buy a company but not an open ecosystem.
  • Tenstorrent's three-legged business model

    1. Hardware: TT-QuietBox 2 workstation ($9,999, 4× Blackhole, runs 120B-parameter models locally), TT-LoudBox servers, PCIe cards, Galaxy Blackhole rack-scale systems. 2. IP licensing (the ARM play): Ascalon RISC-V CPU IP licensed to LG, Hyundai, Japan's LSTC; Neo AI core IP for sovereign AI projects; Open Chiplet Atlas interconnect standard. Most bookings come from IP deals. 3. Sovereign AI: Japan, Korea, Canada and others building NVIDIA-free compute infrastructure — demand NVIDIA cannot serve due to US export controls.

    Funding and ecosystem landscape

  • Series D (Dec 2024): over $1B raised at a $2.6B valuation (Samsung, Bezos Expeditions, Fidelity, Hyundai, LG, XTX Markets, Baillie Gifford); reported $3.2B target valuation by Nov 2025.
  • Challenger landscape: Cerebras (wafer-scale, IPO'd May 2026, ~$56B), Etched (Sohu transformer ASIC, ~$5B), Groq (acquired, $20B), Tenstorrent (~$3.2B), plus AMD ROCm, Google TPU v7 Ironwood, Amazon Trainium/Inferentia, Microsoft Maia.
  • Valuation logic is shifting from "can you build chips" to "can you build an ecosystem."
  • Conclusion

    Can Tenstorrent topple NVIDIA? Not short-term — the software stack is immature, CUDA lock-in persists, and advantages are workload-specific. Possibly long-term: RISC-V datacenter share is projected to reach 25% in 2026, sovereign AI demand is growing, and IP licensing may scale better than chip sales. Keller isn't betting on beating NVIDIA head-on; he's betting the AI compute market will fragment into multiple specialized architectures — where open architecture and the ARM-style licensing model thrive. Groq got a $20B exit ticket; Keller wants Tenstorrent to become the next ARM — infrastructure nobody can avoid, rather than an acquisition target.

    References

  • Tenstorrent: https://tenstorrent.com
  • TT-Forge GitHub: https://github.com/tenstorrent
  • "NVIDIA's $20B Groq Deal" — Yahoo Finance, 2025-12-26
  • "Tenstorrent Vows to 'Crush Everyone'" — WCCFtech, 2026-05-02
  • "RISC-V in 2026: 25% Market Share" — AEStech, 2026-05-03
  • Cerebras IPO: Nasdaq $CBRS, May 2026, ~$56B valuation
  • "Nvidia to license AI chip challenger Groq's tech" — TechCrunch, 2025-12-24

Tags

#nvidia#tenstorrent#jim-keller#risc-v#open-source-chips#ai-inference#groq#semiconductors

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980944