English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ZetaGPT: Making Positional Information Emerge Instead of Bolting It On

Forum topic · ✨步子哥 · 2026-08-11

Summary

ZetaGPT, a reference implementation by Róisín Luo (University of Galway, Ireland; arXiv 2608.09432), explores removing explicit positional encodings like RoPE from language models. Each Transformer block places a causal state-space module (based on Mamba-style selective SSM equations) before self-attention, so recursive state dynamics naturally encode token order—position information 'grows' from the architecture rather than being externally attached. The architecture adds gated multi-head attention to mitigate attention sinks, followed by a standard FFN. Three sub-billion-parameter configurations are released (34.4M, 159.2M, 479.9M params), trained end-to-end on WikiText-103 with byte-level BPE (vocab 50,259), then SFT on Alpaca-GPT4, reward modeling, RLHF, and GRPO-based chain-of-thought training on GSM8K without supervised CoT data. A notable emergent phenomenon: SSM channels spontaneously develop multiple memory time scales during training. The post candidly notes limitations—small scale, narrow pretraining data, and no direct RoPE-vs-no-RoPE comparison—but argues the work's value is proving positional encodings are replaceable, providing the first open-source positional-encoding-free small language model as a clean research baseline.

ZetaGPT: A Reference Implementation of Positional-Encoding-Free State-Space-Attention Language Models

  • Author: Róisín Luo (University of Galway, Ireland)
  • Paper: arXiv: 2608.09432
  • Code: https://github.com/roisincrtai/zetagpt
  • A 'Standard' Design Choice Worth Questioning

    Nearly all mainstream LLMs—from GPT-NeoX to Llama, Qwen, and DeepSeek—attach position information to token representations via RoPE or similar mechanisms. This is because self-attention is permutation-equivariant: it cannot distinguish "the cat chased the dog" from "the dog chased the cat" on its own.

    Positional encodings solve this, but at a cost: when context length exceeds the training range, encodings need extrapolation or fine-tuning. Techniques like NTK-aware scaling, YaRN, and LongRoPE are essentially patches for this problem.

    ZetaGPT asks: can position information emerge from the model's structure rather than being bolted on?

    Core Design: State-Space Module Before Attention

    Each Transformer block has three residual sublayers:

    1. Causal state-space module (SSM): Based on Mamba's selective state-space equations. Recurrent state dynamics encode sequence position into token representations *before* attention sees them. 2. Gated multi-head attention: Standard attention over position-aware representations, with gating to mitigate attention sinks. 3. Feed-forward network: Standard MLP.

    The key is SSM before attention—position comes from recurrent dynamics, not an external encoding.

    Why Recurrence Encodes Position

    Like a reader maintaining a cumulative mental state of a novel, an SSM updates a hidden state:

    \[h_t = a_t \odot h_{t-1} + b_t \odot x_t\]

    where \(a_t\) is an input-dependent decay and \(b_t\) an input-dependent input gate. The state at step \(t\) inherently contains accumulated information from all prior tokens.

    An interesting emergent phenomenon: during training, different SSM channels spontaneously develop multiple memory time scales. With memory horizon \(\tau = -(\ln a_t)^{-1}\), channels diverge within the first 12,800 steps (19.7% of the training budget)—some keeping short \(\tau\) (fast forgetting, local context) and others long \(\tau\) (slow forgetting, long-range dependencies). This mirrors how biological neurons have diverse time constants.

    Model Configurations and Training Pipeline

    | Config | Layers | Dim | Params | |--------|--------|-----|--------| | ZetaGPT-S (default) | 6 | 384 | 34.4M | | ZetaGPT-M | 12 | 768 | 159.2M | | ZetaGPT-L | 24 | 1024 | 479.9M |

    All are sub-billion models aimed at research, prototyping, and education. The end-to-end pipeline:

    1. Tokenizer: Byte-level BPE, vocab size 50,259 (no OOV issues). 2. Pretraining: Language modeling on WikiText-103, 64,840 steps, lr 2×10⁻⁵; loss drops from 10.90 to 6.05 nats/token. 3. SFT: Alpaca-GPT4 instruction data, 2,342 steps. 4. Reward model: Reusing Alpaca-GPT4, 2,500 steps. 5. RLHF: Standard human-feedback reinforcement learning. 6. Chain-of-thought: GRPO on GSM8K—reasoning emerges via pure RL, no supervised CoT data.

    Notably, every stage runs on the same architecture—no component swaps from tokenizer to RLHF to CoT.

    Why 'No Positional Encoding' Matters

  • Context extrapolation: SSM position information accumulates recurrently; architecturally there is no hard length limit (though the paper does not directly test long-context extrapolation).
  • Beyond 'patching': RoPE is an external intervention applied at attention time; ZetaGPT lets position emerge from structure, closer to biological sequence processing.
  • A research baseline: Since virtually all mainstream models use positional encodings, we don't actually know what happens without them. ZetaGPT offers a reproducible baseline for studying that question.
  • Gated Attention as a Bonus

    Attention sinks—attention concentrated on semantically empty tokens like sentence-initial "the"—have been linked to hallucination. ZetaGPT's gated multi-head attention (following Qiu et al. 2025) lets the model dynamically modulate attention contributions, shown to mitigate attention sinks and activation collapse on pre-layer-norm Transformers.

    An Honest Assessment

    Limitations:

  • Small scale: 479.9M max parameters is a toy model by 2026 standards; conclusions may not transfer to billion-parameter models.
  • Narrow pretraining data: Only WikiText-103; no full results on benchmarks like MMLU or HumanEval.
  • No direct RoPE comparison: No same-size, same-data "with vs. without RoPE" experiment, leaving the core question without a clean answer.
  • Single author: Limited experimental coverage; reproducibility needs community verification.
  • Still, the value lies in providing a clean reference implementation—the first open-source positional-encoding-free small language model, with code, data, and training pipeline fully open.

    Part of a 'Change the Level of the Problem' Lineage

    ZetaGPT joins a family of solutions that change the level at which a problem is addressed—RNA editing (blueprint level), slime mold externalized memory (medium level), quantum magnetoreception (physical level), Möbius RoPE (topology level), Euclid-MCP outsourcing reasoning to Prolog (division-of-labor level), and more. ZetaGPT is the eleventh: solving position at the representation level—no bolt-on encoding, just emergent recurrent state dynamics.

    The core insight: many problems that seem to require *addition* (adding encodings, parameters, data) can be solved by *substitution*—changing representation, optimization granularity, or evaluation construction.

    Closing Thoughts

    ZetaGPT won't change the LLM landscape—a 34.4M-parameter model won't compete with frontier systems. But it proves a design choice is replaceable: positional encoding is not sacred, and an entire training pipeline (tokenizer through RLHF to CoT) runs on the alternative. That opens a search space: at what scale does long-context capability emerge? How far can multi-scale memory horizons extend? Can gated attention reduce hallucination at larger scales?

    A good reference implementation is worth more than a benchmark-chasing paper.

    ---

  • Paper: https://arxiv.org/abs/2608.09432
  • Code: https://github.com/roisincrtai/zetagpt

Tags

#state-space-models#positional-encoding#transformers#mamba#rlhf#attention-sink#open-source#language-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633327