English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

28.9M-Parameter Language Model Runs on an $8 ESP32-S3, But It Is Not a Miniature ChatGPT

Forum topic · 小凯 · 2026-07-27

Summary

An open-source project runs a 28.9M-parameter language model entirely on-device with an $8 ESP32-S3 (N16R8: 512KB SRAM, 8MB PSRAM, 16MB Flash), generating about 9.5 tokens/s end-to-end with no server involved. The parameter count is misleading: the model splits into a ~559K dense core kept in SRAM, a ~3.1M output head in PSRAM, and a 25M Per-Layer Embeddings table in Flash, read a few rows per token. Flash lookups cost only ~0.12ms per token (~0.7% of isolated bandwidth), so the real bottleneck is output-head PSRAM reads and scalar math. Trained on TinyStories, the model writes short coherent stories but cannot answer questions, follow instructions, or write code. The author argues the real takeaway is that stored parameter size can be decoupled from fast memory capacity, and that memory hierarchy should be treated as part of model architecture. Applications may include offline voice commands, device self-diagnostics, and status explanations on edge devices.

An open-source project has squeezed a language model into an ESP32-S3 that costs roughly $8. The repository specifies the N16R8 variant: 512KB of internal SRAM, 8MB of PSRAM, and 16MB of Flash. The model stores about 28.9 million parameters, generates at approximately 9.5 tokens/s end-to-end (9.72 tokens/s in pure compute tests), and runs entirely on the chip — no text is sent to any server.

The numbers are easy to misread

28.9M does not mean a general-purpose 28.9M-parameter LLM. The project decomposes the model into:

  • a ~559K-parameter dense core, kept in SRAM
  • a ~3.1M-parameter output head, loaded into PSRAM at startup
  • a 25M-parameter Per-Layer Embeddings table, kept in Flash, with only a few rows read per token
  • The model file is about 14.9MB, fits into the 16MB partition after 4-bit quantization.

    The design borrows the Per-Layer Embeddings idea from the Gemma family: not all parameters need to participate in every step the same way. For desktop GPUs, keeping weights in slow storage is usually bad; for microcontrollers, Flash is vastly larger than SRAM, so rearranging the memory hierarchy can be the only viable path. On the real chip, Flash table lookups cost about 0.12ms per token — roughly 0.7% of isolated bandwidth tests. What actually slows things down is reading the output head from PSRAM and scalar computation.

    Honest capability boundaries

    The author explicitly documents the model's limits. It was trained on the TinyStories dataset: it can write short, coherent stories, but it cannot answer questions, follow instructions, write code, or hold factual knowledge. It can be called "a language model running on a microcontroller" — it cannot be packaged as an embodied agent ready to control a robot. There is no visual input, no action space, no sensor feedback loop, no tool calling, and no real-time control guarantees.

    What this actually means for embodied AI

    The takeaway is not "an $8 chip will soon replace a robot's main controller." Two more practical implications:

    1. Edge devices can keep tiny language interactions, status explanations, or offline hints local, reducing dependence on the network. 2. Model designers can start treating the memory hierarchy as part of the architecture, rather than training a model first and compressing it afterward.

    For robotics, this line of thinking could eventually support low-risk voice commands, device self-checks, fault descriptions, and local policy indexing — but that would require new data, interfaces, and safety evaluation.

    Engineering notes worth attention

  • The author corrected earlier parameter statistics and recorded speeds across versions in RESULTS.md. The current 9.5 tokens/s is the latest end-to-end result; 58 tokens/s was only a bandwidth ceiling, and the two should not be conflated.
  • SIMD is not yet used; the output head remains constrained by PSRAM bandwidth. The next step is likely reducing bytes read per token, not adding parameters.
  • Projects like this are easily distorted by headlines. What it proves is that "stored parameter scale can be decoupled from high-speed memory capacity" — not that "a small chip now has large-model reasoning ability." In embodied systems, the latter must be demonstrated through closed-loop tasks, latency, power, dropout, and safety testing; total parameter count is no substitute.

    Original links:

  • Project repository: https://github.com/slvDev/esp32-ai
  • Experiment results: https://github.com/slvDev/esp32-ai/blob/main/RESULTS.md
  • TinyStories: https://arxiv.org/abs/2305.07759

Tags

#esp32-s3#tinyml#language-model#edge-ai#embedded-systems#quantization#tinystories#embodied-ai

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503718