Jim Keller's Tenstorrent Bet: Can Open-Source RISC-V Chips Topple NVIDIA's Empire?
> A note on the source: a video title misspells the company as "TenStorent" — itself a sign that Tenstorrent's chips aren't yet big enough on the table for everyone to know its name. But Jim Keller's résumé commands attention: AMD Zen, Apple A4/A5, Tesla FSD, Intel Xe. The 64-year-old chip legend is now making the last big bet of his career: using open-source RISC-V architecture to take AI inference market share from NVIDIA. The stakes are real — in December 2025, NVIDIA reportedly spent ~$20 billion on Groq's LPU technology and team, a move the author frames as an "execution" rather than an acquisition. Does Tenstorrent have something NVIDIA can't buy?
Key points
- Jim Keller's track record: AMD K7/K8 (first real threat to Intel in x86), Apple A4/A5 (foundation of the iPhone/iPad), AMD Zen (rescued AMD), Tesla FSD (first automotive AI inference chip), Intel Xe (unsuccessful), now Tenstorrent.
- Philosophy: don't optimize what others have optimized — redefine the problem. Tenstorrent is his hardest hand yet against a $3-trillion, full-stack NVIDIA.
- Each Tensix core contains 5 RISC-V processors plus vector/matrix units and local SRAM.
- No hardware cache hierarchy (L1/L2/L3); data movement is explicitly software-controlled.
- Chips connect via on-chip Ethernet in a mesh/torus topology — no NVLink or external switches required.
- TT-Forge: open AI compiler supporting PyTorch/JAX/ONNX (public beta)
- TT-Metalium: low-level SDK for writing kernels, MIT license
- TT-LLK: low-level kernel software
- RISC-V ISA: royalty-free open instruction set
- Galaxy Blackhole server: DeepSeek R1 inference at 350+ tokens/second
- TCO 5x lower than NVIDIA GB300
- "Blitz Mode": 10x faster generative AI video vs. current GPUs
- Keller: "We are committed to crushing everybody at everything"
- Groq's LPU (designed by Google TPU father Jonathan Ross) removes caches, branch prediction, and OoO execution entirely, using compiler-deterministic execution and on-chip SRAM — hitting 241 tokens/s on Llama 2 70B in early 2024.
- On December 24, 2025, NVIDIA announced a ~$20 billion deal for Groq's core technology and team (including Ross and Sunny Madra) — nearly 3x the Mellanox deal.
- Motives: GPUs are weak at the memory-bandwidth-bound decode phase where Groq's SRAM architecture excels; buying Groq removes a threat and absorbs key talent.
- Lesson for Tenstorrent: novel inference architectures have value, but the ceiling may be an acquisition price. Keller's counter is open source — NVIDIA can buy a company but not an open ecosystem.
- Series D (Dec 2024): over $1B raised at a $2.6B valuation (Samsung, Bezos Expeditions, Fidelity, Hyundai, LG, XTX Markets, Baillie Gifford); reported $3.2B target valuation by Nov 2025.
- Challenger landscape: Cerebras (wafer-scale, IPO'd May 2026, ~$56B), Etched (Sohu transformer ASIC, ~$5B), Groq (acquired, $20B), Tenstorrent (~$3.2B), plus AMD ROCm, Google TPU v7 Ironwood, Amazon Trainium/Inferentia, Microsoft Maia.
- Valuation logic is shifting from "can you build chips" to "can you build an ecosystem."
- Tenstorrent: https://tenstorrent.com
- TT-Forge GitHub: https://github.com/tenstorrent
- "NVIDIA's $20B Groq Deal" — Yahoo Finance, 2025-12-26
- "Tenstorrent Vows to 'Crush Everyone'" — WCCFtech, 2026-05-02
- "RISC-V in 2026: 25% Market Share" — AEStech, 2026-05-03
- Cerebras IPO: Nasdaq $CBRS, May 2026, ~$56B valuation
- "Nvidia to license AI chip challenger Groq's tech" — TechCrunch, 2025-12-24
Tenstorrent's technical approach: neither GPU nor TPU
Tensix cores
| Dimension | NVIDIA GPU | Tenstorrent Tensix | |---|---|---| | Control | SIMT | 5 independent RISC-V cores | | Memory | Hardware-managed caches | Explicit SRAM management, no caches | | Interconnect | NVLink/InfiniBand + switches | On-chip Ethernet, direct torus | | Software | CUDA (proprietary) | TT-Forge/TT-Metalium (open source) | | Scaling | Dedicated interconnect hardware | Standard Ethernet |
Keller's bet: AI dataflow patterns are predictable, so transistors spent on caches, branch prediction, and out-of-order execution are wasted — better to spend them on matrix compute and let software move data precisely.
Fully open-source software stack
The strategy: replicate Linux's victory over Unix — use open-ecosystem long-term stickiness to counter CUDA's proprietary short-term advantages.
Chip roadmap
| Generation | Timing | Notes | |---|---|---| | Grayskull | 2020 | Proof of concept | | Wormhole | 2022–2023 | Better interconnect, training + inference | | Blackhole | 2024–2026 | Current flagship: 480 Tensix cores, 2,654 TFLOPS (BlockFP8) | | Grendel | 2026+ | Chiplet architecture (Open Chiplet Atlas), separate CPU and AI tiles |Performance claims vs. reality
Tenstorrent's May 2026 TT-Deploy event claimed:
The article urges skepticism: undisclosed test conditions (batch size, quantization, prefill/decode split) make token-rate claims hard to verify; TCO figures may ignore software migration and retraining costs; the "Blitz Mode" comparison baseline is unclear; and TT-Forge's beta status means production readiness remains unproven.
Galaxy Blackhole vs. GB300
| Dimension | NVIDIA GB300 | Galaxy Blackhole | |---|---|---| | Memory | HBM3e | GDDR6 / on-chip SRAM | | Ecosystem | CUDA (20 years) | TT-Forge (beta) | | Cloud availability | All major clouds | Koyeb and partners | | Model support | Nearly all AI models | Compile-dependent, limited | | Production validation | Massive global deployment | Small deployments, Japan/Korea sovereign projects |
Bottom line: Tenstorrent may win on specific workloads, but generality and ecosystem maturity trail NVIDIA by orders of magnitude — a 5-year startup vs. a 20-year incumbent.
The $20B Groq precedent
Tenstorrent's three-legged business model
1. Hardware: TT-QuietBox 2 workstation ($9,999, 4× Blackhole, runs 120B-parameter models locally), TT-LoudBox servers, PCIe cards, Galaxy Blackhole rack-scale systems. 2. IP licensing (the ARM play): Ascalon RISC-V CPU IP licensed to LG, Hyundai, Japan's LSTC; Neo AI core IP for sovereign AI projects; Open Chiplet Atlas interconnect standard. Most bookings come from IP deals. 3. Sovereign AI: Japan, Korea, Canada and others building NVIDIA-free compute infrastructure — demand NVIDIA cannot serve due to US export controls.
Funding and ecosystem landscape
Conclusion
Can Tenstorrent topple NVIDIA? Not short-term — the software stack is immature, CUDA lock-in persists, and advantages are workload-specific. Possibly long-term: RISC-V datacenter share is projected to reach 25% in 2026, sovereign AI demand is growing, and IP licensing may scale better than chip sales. Keller isn't betting on beating NVIDIA head-on; he's betting the AI compute market will fragment into multiple specialized architectures — where open architecture and the ARM-style licensing model thrive. Groq got a $20B exit ticket; Keller wants Tenstorrent to become the next ARM — infrastructure nobody can avoid, rather than an acquisition target.
References