English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Kronos: The First Open-Source Foundation Model for Financial Markets — Treating Candlesticks as Language

Forum topic · ✨步子哥 · 2026-08-03

Summary

Kronos is the first open-source foundation model pretrained specifically for financial market time series, accepted at AAAI 2026 (arXiv: 2508.02739). Its key insight is treating OHLCV candlestick data as a language: a specialized hierarchical tokenizer quantizes continuous open/high/low/close/volume data into discrete tokens at multiple scales, and an autoregressive Transformer — the same recipe as GPT — is pretrained on these token sequences. The model family includes Kronos-mini (4.1M parameters, 2048 context), Kronos-small (24.7M), and Kronos-base (102.3M), all open-sourced; Kronos-large (499.2M) remains unreleased. Trained on candlestick data from 45+ global exchanges, Kronos addresses the failure of general time series foundation models on financial data, which suffers from high noise, non-stationarity, and fat-tailed distributions. Hierarchical quantization preserves multi-scale price movements, mirroring how human analysts read daily, weekly, and monthly charts simultaneously. The post frames Kronos as part of a broader "X as language" paradigm: discretize any raw sequential data into tokens, then apply Transformer pretraining.

Kronos: Treating Candlesticks as a Language — The First Open-Source Foundation Model for Financial Markets

Technical analysts read candlestick charts and see head-and-shoulders tops, double bottoms, hammer lines — "patterns" that are essentially vocabulary the human eye identifies in price series.

What would an AI see if it looked at candlesticks?

Kronos's answer: treat candlesticks as a language, and use language-model methods to understand them.

The Old Problem of Financial Time Series Forecasting

Financial time series forecasting is notoriously difficult. Traditional methods (ARIMA, LSTM) treat prices as continuous numbers, but financial data has three fatal characteristics:

1. High noise: Most price fluctuations are noise, not signal 2. Non-stationarity: Distributions drift over time; patterns that work today fail tomorrow 3. Fat-tailed distributions: Extreme events occur more frequently than normal distributions predict

General-purpose time series foundation models (TSFMs) underperform on finance because they never learned the "dialect" of financial data. Trained on relatively clean series like electricity load or traffic flow, they struggle with the violent noise of financial markets.

Kronos's Two-Stage Framework

Kronos (GitHub Trending +217⭐/day, accepted at AAAI 2026) is the first open-source foundation model pretrained specifically for financial markets. Its core innovation is a two-stage framework:

Stage 1: A Specialized Tokenizer Quantizes OHLCV into Hierarchical Discrete Tokens

OHLCV (open/high/low/close/volume) is continuous multi-dimensional data. Rather than simple uniform bucketing, Kronos uses hierarchical quantization — capturing price movements at different scales.

This step is critical. Uniform bucketing loses multi-scale features of price movement — a 1% move and a 5% move might land in the same bucket. Hierarchical quantization gives small and large fluctuations their own "vocabulary."

Stage 2: Autoregressive Transformer Pretraining

Autoregressive pretraining on the hierarchical discrete tokens, just like GPT — except the "vocabulary" is candlestick tokens instead of natural language.

The elegance of this framework: it didn't invent a new architecture; it invented a new "language." The Transformer and pretraining recipe are off-the-shelf; the innovation lies in how continuous financial data gets "translated" into discrete tokens.

Data and Models

Training Data

Pretrained on candlestick data from 45+ global exchanges. This coverage means Kronos learned not just the "dialect" of US or Chinese markets, but the "lingua franca" of global markets.

Model Family

| Model | Parameters | Context Length | Open Source | |---|---|---|---| | Kronos-mini | 4.1M | 2048 | ✅ | | Kronos-small | 24.7M | 512 | ✅ | | Kronos-base | 102.3M | 512 | ✅ | | Kronos-large | 499.2M | 512 | ❌ |

Note that Kronos-mini has only 4.1M parameters with 2k context — remarkably small. GPT-class models run into tens or hundreds of billions of parameters, yet Kronos-mini can forecast with 4.1M. This suggests the "language" of financial time series is far simpler than natural language — small vocabulary, few grammar rules, short context dependencies.

Academic Credentials

Accepted at AAAI 2026, arXiv: 2508.02739, with full comparative experiments and ablation studies in the paper.

Engineering Insight: X as Language

Kronos points to a deeper trend — the "X as language" framework is expanding into more and more domains:

  • Natural language GPT: text sequences as language
  • VLMs: image patches discretized into tokens
  • Speech models: sound waves discretized into tokens
  • Kronos: candlesticks discretized into tokens
  • The common thread: "discretize continuous/raw data into tokens, then pretrain a Transformer." The real power of Transformers isn't in natural language — it's in any sequence that can be discretized into tokens.

    The Subtlety of Hierarchical Quantization

    "Hierarchical quantization" sounds simple, but it is Kronos's core innovation. An analogy:

  • Uniform bucketing = treating all price changes at equal granularity — if 1% and 5% moves land in the same bucket, scale information is lost
  • Hierarchical quantization = multiple granularities coexist — coarse-grained for trends, fine-grained for fluctuations, each scale with its own vocabulary
  • This is isomorphic to how humans read charts. Analysts don't look only at 1-minute candles; they simultaneously study daily, weekly, and monthly charts — different scales reveal different signals. Kronos's hierarchical quantization encodes this multi-scale perspective into the tokenizer.

    Cross-Paper Consensus

    Kronos's design resonates with recent work:

  • Natively multimodal Scaling Laws: different modalities have different "languages"; specialized tokenizers are key — Kronos's financial tokenizer confirms this
  • Möbius RoPE: positional encoding matters for time series — Kronos's context length choices (mini 2k, others 512) reflect consideration of temporal dependency depth
  • Octopus RNA editing: don't change the blueprint (raw OHLCV data); change the construction plan (tokenization approach) — Kronos's innovation is in data representation, not model architecture

Conclusion

The trend Kronos suggests is bigger than financial forecasting itself — the "X as language" framework is expanding. Financial candlesticks are just the beginning; "protein languages," "climate languages," and "genome languages" may follow. Once any continuous data can be discretized into tokens, the Transformer's application boundary extends from natural language to all sequential data.

The first open-source foundation model for financial markets isn't just another tool for quantitative traders — it's another validation of the "X as language" paradigm.

---

Project: https://github.com/shiyu-coder/Kronos Paper: https://arxiv.org/abs/2508.02739 Live Demo: https://shiyu-coder.github.io/Kronos-demo/ Models: HuggingFace NeoQuasar/Kronos-{mini,small,base} License: Open source Academic: Accepted at AAAI 2026

Tags

#kronos#foundation-models#financial-time-series#transformers#tokenization#quantitative-finance#autoregressive-models#aaai-2026

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503924