ZetaGPT: A Reference Implementation of Positional-Encoding-Free State-Space-Attention Language Models
- Author: Róisín Luo (University of Galway, Ireland)
- Paper: arXiv: 2608.09432
- Code: https://github.com/roisincrtai/zetagpt
- Context extrapolation: SSM position information accumulates recurrently; architecturally there is no hard length limit (though the paper does not directly test long-context extrapolation).
- Beyond 'patching': RoPE is an external intervention applied at attention time; ZetaGPT lets position emerge from structure, closer to biological sequence processing.
- A research baseline: Since virtually all mainstream models use positional encodings, we don't actually know what happens without them. ZetaGPT offers a reproducible baseline for studying that question.
- Small scale: 479.9M max parameters is a toy model by 2026 standards; conclusions may not transfer to billion-parameter models.
- Narrow pretraining data: Only WikiText-103; no full results on benchmarks like MMLU or HumanEval.
- No direct RoPE comparison: No same-size, same-data "with vs. without RoPE" experiment, leaving the core question without a clean answer.
- Single author: Limited experimental coverage; reproducibility needs community verification.
- Paper: https://arxiv.org/abs/2608.09432
- Code: https://github.com/roisincrtai/zetagpt
A 'Standard' Design Choice Worth Questioning
Nearly all mainstream LLMs—from GPT-NeoX to Llama, Qwen, and DeepSeek—attach position information to token representations via RoPE or similar mechanisms. This is because self-attention is permutation-equivariant: it cannot distinguish "the cat chased the dog" from "the dog chased the cat" on its own.
Positional encodings solve this, but at a cost: when context length exceeds the training range, encodings need extrapolation or fine-tuning. Techniques like NTK-aware scaling, YaRN, and LongRoPE are essentially patches for this problem.
ZetaGPT asks: can position information emerge from the model's structure rather than being bolted on?
Core Design: State-Space Module Before Attention
Each Transformer block has three residual sublayers:
1. Causal state-space module (SSM): Based on Mamba's selective state-space equations. Recurrent state dynamics encode sequence position into token representations *before* attention sees them. 2. Gated multi-head attention: Standard attention over position-aware representations, with gating to mitigate attention sinks. 3. Feed-forward network: Standard MLP.
The key is SSM before attention—position comes from recurrent dynamics, not an external encoding.
Why Recurrence Encodes Position
Like a reader maintaining a cumulative mental state of a novel, an SSM updates a hidden state:
where \(a_t\) is an input-dependent decay and \(b_t\) an input-dependent input gate. The state at step \(t\) inherently contains accumulated information from all prior tokens.
An interesting emergent phenomenon: during training, different SSM channels spontaneously develop multiple memory time scales. With memory horizon \(\tau = -(\ln a_t)^{-1}\), channels diverge within the first 12,800 steps (19.7% of the training budget)—some keeping short \(\tau\) (fast forgetting, local context) and others long \(\tau\) (slow forgetting, long-range dependencies). This mirrors how biological neurons have diverse time constants.
Model Configurations and Training Pipeline
| Config | Layers | Dim | Params | |--------|--------|-----|--------| | ZetaGPT-S (default) | 6 | 384 | 34.4M | | ZetaGPT-M | 12 | 768 | 159.2M | | ZetaGPT-L | 24 | 1024 | 479.9M |
All are sub-billion models aimed at research, prototyping, and education. The end-to-end pipeline:
1. Tokenizer: Byte-level BPE, vocab size 50,259 (no OOV issues). 2. Pretraining: Language modeling on WikiText-103, 64,840 steps, lr 2×10⁻⁵; loss drops from 10.90 to 6.05 nats/token. 3. SFT: Alpaca-GPT4 instruction data, 2,342 steps. 4. Reward model: Reusing Alpaca-GPT4, 2,500 steps. 5. RLHF: Standard human-feedback reinforcement learning. 6. Chain-of-thought: GRPO on GSM8K—reasoning emerges via pure RL, no supervised CoT data.
Notably, every stage runs on the same architecture—no component swaps from tokenizer to RLHF to CoT.
Why 'No Positional Encoding' Matters
Gated Attention as a Bonus
Attention sinks—attention concentrated on semantically empty tokens like sentence-initial "the"—have been linked to hallucination. ZetaGPT's gated multi-head attention (following Qiu et al. 2025) lets the model dynamically modulate attention contributions, shown to mitigate attention sinks and activation collapse on pre-layer-norm Transformers.
An Honest Assessment
Limitations:
Still, the value lies in providing a clean reference implementation—the first open-source positional-encoding-free small language model, with code, data, and training pipeline fully open.
Part of a 'Change the Level of the Problem' Lineage
ZetaGPT joins a family of solutions that change the level at which a problem is addressed—RNA editing (blueprint level), slime mold externalized memory (medium level), quantum magnetoreception (physical level), Möbius RoPE (topology level), Euclid-MCP outsourcing reasoning to Prolog (division-of-labor level), and more. ZetaGPT is the eleventh: solving position at the representation level—no bolt-on encoding, just emergent recurrent state dynamics.
The core insight: many problems that seem to require *addition* (adding encodings, parameters, data) can be solved by *substitution*—changing representation, optimization granularity, or evaluation construction.
Closing Thoughts
ZetaGPT won't change the LLM landscape—a 34.4M-parameter model won't compete with frontier systems. But it proves a design choice is replaceable: positional encoding is not sacred, and an entire training pipeline (tokenizer through RLHF to CoT) runs on the alternative. That opens a search space: at what scale does long-context capability emerge? How far can multi-scale memory horizons extend? Can gated attention reduce hallucination at larger scales?
A good reference implementation is worth more than a benchmark-chasing paper.
---