An open-source project has squeezed a language model into an ESP32-S3 that costs roughly $8. The repository specifies the N16R8 variant: 512KB of internal SRAM, 8MB of PSRAM, and 16MB of Flash. The model stores about 28.9 million parameters, generates at approximately 9.5 tokens/s end-to-end (9.72 tokens/s in pure compute tests), and runs entirely on the chip — no text is sent to any server.
The numbers are easy to misread
28.9M does not mean a general-purpose 28.9M-parameter LLM. The project decomposes the model into:
- a ~559K-parameter dense core, kept in SRAM
- a ~3.1M-parameter output head, loaded into PSRAM at startup
- a 25M-parameter Per-Layer Embeddings table, kept in Flash, with only a few rows read per token
- The author corrected earlier parameter statistics and recorded speeds across versions in RESULTS.md. The current 9.5 tokens/s is the latest end-to-end result; 58 tokens/s was only a bandwidth ceiling, and the two should not be conflated.
- SIMD is not yet used; the output head remains constrained by PSRAM bandwidth. The next step is likely reducing bytes read per token, not adding parameters.
- Project repository: https://github.com/slvDev/esp32-ai
- Experiment results: https://github.com/slvDev/esp32-ai/blob/main/RESULTS.md
- TinyStories: https://arxiv.org/abs/2305.07759
The model file is about 14.9MB, fits into the 16MB partition after 4-bit quantization.
The design borrows the Per-Layer Embeddings idea from the Gemma family: not all parameters need to participate in every step the same way. For desktop GPUs, keeping weights in slow storage is usually bad; for microcontrollers, Flash is vastly larger than SRAM, so rearranging the memory hierarchy can be the only viable path. On the real chip, Flash table lookups cost about 0.12ms per token — roughly 0.7% of isolated bandwidth tests. What actually slows things down is reading the output head from PSRAM and scalar computation.
Honest capability boundaries
The author explicitly documents the model's limits. It was trained on the TinyStories dataset: it can write short, coherent stories, but it cannot answer questions, follow instructions, write code, or hold factual knowledge. It can be called "a language model running on a microcontroller" — it cannot be packaged as an embodied agent ready to control a robot. There is no visual input, no action space, no sensor feedback loop, no tool calling, and no real-time control guarantees.
What this actually means for embodied AI
The takeaway is not "an $8 chip will soon replace a robot's main controller." Two more practical implications:
1. Edge devices can keep tiny language interactions, status explanations, or offline hints local, reducing dependence on the network. 2. Model designers can start treating the memory hierarchy as part of the architecture, rather than training a model first and compressing it afterward.
For robotics, this line of thinking could eventually support low-risk voice commands, device self-checks, fault descriptions, and local policy indexing — but that would require new data, interfaces, and safety evaluation.
Engineering notes worth attention
Projects like this are easily distorted by headlines. What it proves is that "stored parameter scale can be decoupled from high-speed memory capacity" — not that "a small chip now has large-model reasoning ability." In embodied systems, the latter must be demonstrated through closed-loop tasks, latency, power, dropout, and safety testing; total parameter count is no substitute.
Original links: