English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

55nm ReRAM-on-Logic Stacked Chip Runs LLM Inference at up to 135 tokens/s (ISSCC 2026)

Forum topic · 小凯 · 2026-05-18

Summary

A forum post discusses an ISSCC 2026 LLM inference accelerator built on a 55nm process using bump-bonded face-to-face ReRAM-on-Logic vertical stacking, reporting throughput of 14.08 to 135.69 tokens/s (arXiv:2605.09375). Instead of HBM, the design uses ReRAM (resistive random-access memory), which offers higher density than SRAM, lower power than DRAM, and the ability to stack directly on the logic layer. Key techniques include outlier-free low-bit quantization via local rotating units, block-clustering vector compression to reduce weight-loading overhead, and adaptive parallel speculative decoding. The author raises caveats: 55nm is an older node, so comparisons with accelerators on advanced processes may not be fair, and the wide token/s range leaves unclear which configurations produce 14 versus 135 tokens/s.

LLM accelerators typically rely on HBM for high-capacity, high-bandwidth memory. ReRAM (resistive random-access memory) is an alternative: it has higher density than SRAM, lower power than DRAM, and can be vertically stacked with logic layers.

A chip presented at ISSCC 2026 (arXiv:2605.09375) implements an LLM inference accelerator in 55nm CMOS using bump-bonded face-to-face ReRAM-on-Logic stacking, delivering 14.08–135.69 token/s. Core techniques include:

  • Outlier-free low-bit quantization using local rotating units
  • Block-clustering vector compression to reduce weight-loading overhead
  • Adaptive parallel speculative decoding
  • Open questions raised by the author

  • 55nm is a relatively old process node. When comparing against accelerators in the literature built on more advanced nodes, fairness requires accounting for process differences.
  • The 14–135 token/s range is very wide — under what conditions does the chip run at 14 token/s versus 135 token/s?

References

1. Dong, P., et al. (2026). *A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked LLM Accelerator*. arXiv:2605.09375 [cs.AR]. (ISSCC 2026) 2. Chen, A. (2024). *ReRAM-based Processing-in-Memory for AI*. Nature Electronics. 3. Levi, T., et al. (2023). *Speculative Decoding for LLM Inference Acceleration*.

Tags

#llm-inference#reram#isscc-2026#accelerators#processing-in-memory#speculative-decoding#quantization#chip-stacking

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620287