English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

A 600-Year-Old Unreadable Book: Statistician Confirms the Voynich Manuscript Contains Two Distinct 'Languages'

Forum topic · 二一 · 2026-05-01

Summary

Christophe Parisel's 2026 paper (arXiv:2604.25979) provides quantitative confirmation of Prescott Currier's 1976 hypothesis that the Voynich Manuscript contains two distinct statistical systems, known as Currier A and Currier B. Using Beta-Binomial mixture models, Parisel shows that an unsupervised clustering algorithm, given no prior labels, independently recovers a two-regime structure (k=2, ARI=0.383 versus Currier's hand labels), with 61% of folios assigned at over 90% posterior confidence. Label-permutation tests, single-source Markov simulations, and cross-validated prediction (89.2% accuracy) all support the distinction. Character pairs fall into three functional zones: categorical, intermediate, and free variation. However, the A/B split explains only 29.3% of variance, and the asymmetry of the signal weakens hoax theories based on generated gibberish. The results impose new constraints on any future decipherment attempt.

> Paper: A Quantitative Confirmation of the Currier Language Distinction > Author: Christophe Parisel > arXiv: 2604.25979 [cs.CR, cs.CL] > Published: 2026-04-28

This post discusses a new statistical study of the Voynich Manuscript — the famous ~240-page early-15th-century (carbon-dated 1404–1438) illuminated codex discovered by book dealer Wilfrid Voynich in 1912, whose unknown script has resisted decipherment for six centuries despite efforts from Enigma codebreakers, NSA researchers, and modern machine learning.

Key points

Currier's 1976 hypothesis

  • NSA cryptologist Prescott Currier proposed in 1976 that the manuscript's text falls into two statistically distinct "languages," later called Currier A and Currier B, based on token frequencies (e.g., "chol"/"chor" common in A; word-final "dy" common in B; "chedy" absent from A).
  • Palaeographer Lisa Fagin Davis later identified five scribal hands — but no hand writes both A and B (Scribes 1 and 4 write A; Scribes 2, 3, and 5 write B).
  • Three challenges addressed by Parisel

    1. Could the A/B differences be random fluctuation within a single text? 2. Could they be an artifact of quire (physical booklet) boundaries rather than the text itself? 3. Could an unbiased algorithm discover the distinction without knowing Currier's labels?

    Methods and findings

  • Label shuffling: Randomly permuting A/B labels 1,000 times almost never reproduces the observed pattern (p < 0.001).
  • Single-source Markov simulation: 200 synthetic "manuscripts" generated from one uniform statistical source all fail to replicate the A/B pattern.
  • Unsupervised clustering: A Beta-Binomial mixture model, given only raw character counts from 185 folios, selects k = 2 regimes by BIC, recovering Currier's split with ARI = 0.383; 113 of 185 folios (61%) are assigned at >90% posterior confidence.
  • Predictive validation: A supervised classifier achieves 89.2% cross-validated accuracy on held-out folios; zero of 500 label permutations reach comparable accuracy (p < 0.002).
  • Three functional zones of character pairs

  • Categorical (d/l, or/ar, s/r, e/ee): Cramér's V > 0.20; near-binary switching between A and B — behavior strongly reminiscent of cipher substitution.
  • Intermediate (ol/al, y/dy, k/t): Cramér's V 0.04–0.15; graded preferences rather than absolute switches.
  • Free variation (o/a, ch/sh, f/p): Cramér's V near zero; unaffected by the A/B split, like ordinary allographic variation.
  • The e/ch anomaly

    The e/ch pair has near-zero aggregate association (Cramér's V = 0.007) yet shows the largest jump at A/B boundaries (p = 2.45 × 10⁻⁶). Including it suppresses clustering (K-means ARI 0.208 vs 0.456 without it), suggesting its variation is conditioned by finer-grained factors such as word-level structure or character position.

    Limits of the A/B framework

  • A/B labels explain only 29.3% of variance (R² = 0.293) in character-pair ratios across folios; over 70% of statistical structure remains unexplained.
  • Split-Markov models underestimate per-folio variance by 1.5–6×, implying substructure within A and B.
  • The signal is temporally asymmetric: forward spatial prediction (earlier folios → later) reaches 79.6% accuracy, backward only 41.3%, suggesting A dominates early folios and B later ones — a possible compositional order.
  • Implications

  • Against hoax theories: Gordon Rugg's Cardan-grille hypothesis (2004) and Schinner & Timm's generative algorithms (2020) imply a uniform statistical structure. Parisel's results — no single source can produce the A/B differences (200 failed simulations), and both regimes coexist within single quires (p = 6.89 × 10⁻⁵) — make pure gibberish generation much harder to sustain. A simpler explanation is that the manuscript contains real, meaningful structure.
  • Relation to the Naibbe cipher: Michael Greshko's Cryptologia (2025) construction shows a 15th-century-feasible verbose homophonic cipher can mimic many Voynichese statistics — but cannot fully reproduce Voynich B's properties. Parisel's findings add constraints any candidate cipher must satisfy.
The A/B distinction is real, predictable, and the manuscript's dominant statistical axis — yet it accounts for less than a third of the structure. As the post concludes, the Voynich Manuscript is "a castle whose doors we know exist but whose keyholes we haven't found."

*Based on Christophe Parisel's paper arXiv:2604.25979. Historical background draws on Mary D'Imperio's *An Elegant Enigma* (NSA, 1978), Kennedy & Churchill's *The Voynich Manuscript* (2004), and Bowern & Lindemann's *The Linguistics of the Voynich Manuscript* (Annual Review of Linguistics, 2021). The Naibbe cipher discussion references Michael Greshko's Cryptologia (2025) research.*

Tags

#voynich-manuscript#cryptography#statistics#bayesian-models#unsupervised-learning#currier-languages#medieval-manuscripts#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618961