> Paper: A Quantitative Confirmation of the Currier Language Distinction > Author: Christophe Parisel > arXiv: 2604.25979 [cs.CR, cs.CL] > Published: 2026-04-28
This post discusses a new statistical study of the Voynich Manuscript — the famous ~240-page early-15th-century (carbon-dated 1404–1438) illuminated codex discovered by book dealer Wilfrid Voynich in 1912, whose unknown script has resisted decipherment for six centuries despite efforts from Enigma codebreakers, NSA researchers, and modern machine learning.
Key points
Currier's 1976 hypothesis
- NSA cryptologist Prescott Currier proposed in 1976 that the manuscript's text falls into two statistically distinct "languages," later called Currier A and Currier B, based on token frequencies (e.g., "chol"/"chor" common in A; word-final "dy" common in B; "chedy" absent from A).
- Palaeographer Lisa Fagin Davis later identified five scribal hands — but no hand writes both A and B (Scribes 1 and 4 write A; Scribes 2, 3, and 5 write B).
- Label shuffling: Randomly permuting A/B labels 1,000 times almost never reproduces the observed pattern (p < 0.001).
- Single-source Markov simulation: 200 synthetic "manuscripts" generated from one uniform statistical source all fail to replicate the A/B pattern.
- Unsupervised clustering: A Beta-Binomial mixture model, given only raw character counts from 185 folios, selects k = 2 regimes by BIC, recovering Currier's split with ARI = 0.383; 113 of 185 folios (61%) are assigned at >90% posterior confidence.
- Predictive validation: A supervised classifier achieves 89.2% cross-validated accuracy on held-out folios; zero of 500 label permutations reach comparable accuracy (p < 0.002).
- Categorical (d/l, or/ar, s/r, e/ee): Cramér's V > 0.20; near-binary switching between A and B — behavior strongly reminiscent of cipher substitution.
- Intermediate (ol/al, y/dy, k/t): Cramér's V 0.04–0.15; graded preferences rather than absolute switches.
- Free variation (o/a, ch/sh, f/p): Cramér's V near zero; unaffected by the A/B split, like ordinary allographic variation.
- A/B labels explain only 29.3% of variance (R² = 0.293) in character-pair ratios across folios; over 70% of statistical structure remains unexplained.
- Split-Markov models underestimate per-folio variance by 1.5–6×, implying substructure within A and B.
- The signal is temporally asymmetric: forward spatial prediction (earlier folios → later) reaches 79.6% accuracy, backward only 41.3%, suggesting A dominates early folios and B later ones — a possible compositional order.
- Against hoax theories: Gordon Rugg's Cardan-grille hypothesis (2004) and Schinner & Timm's generative algorithms (2020) imply a uniform statistical structure. Parisel's results — no single source can produce the A/B differences (200 failed simulations), and both regimes coexist within single quires (p = 6.89 × 10⁻⁵) — make pure gibberish generation much harder to sustain. A simpler explanation is that the manuscript contains real, meaningful structure.
- Relation to the Naibbe cipher: Michael Greshko's Cryptologia (2025) construction shows a 15th-century-feasible verbose homophonic cipher can mimic many Voynichese statistics — but cannot fully reproduce Voynich B's properties. Parisel's findings add constraints any candidate cipher must satisfy.
Three challenges addressed by Parisel
1. Could the A/B differences be random fluctuation within a single text? 2. Could they be an artifact of quire (physical booklet) boundaries rather than the text itself? 3. Could an unbiased algorithm discover the distinction without knowing Currier's labels?Methods and findings
Three functional zones of character pairs
The e/ch anomaly
The e/ch pair has near-zero aggregate association (Cramér's V = 0.007) yet shows the largest jump at A/B boundaries (p = 2.45 × 10⁻⁶). Including it suppresses clustering (K-means ARI 0.208 vs 0.456 without it), suggesting its variation is conditioned by finer-grained factors such as word-level structure or character position.Limits of the A/B framework
Implications
*Based on Christophe Parisel's paper arXiv:2604.25979. Historical background draws on Mary D'Imperio's *An Elegant Enigma* (NSA, 1978), Kennedy & Churchill's *The Voynich Manuscript* (2004), and Bowern & Lindemann's *The Linguistics of the Voynich Manuscript* (Annual Review of Linguistics, 2021). The Naibbe cipher discussion references Michael Greshko's Cryptologia (2025) research.*