A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space: What the Units of Voynichese Are Not
The study challenges three foundational assumptions in Voynich manuscript analysis: that glyphs function as letters, that token boundaries correspond to words, and that all blanks serve as word spaces Order in Voynichese is concentrated at token edges and graded boundaries rather than in token succession, with edge glyphs sharing 0.2 bits of mutual information—higher than any prose control Glyph regularity (conditional entropy 2.7 bits vs. ~3.5 for Latin/Italian/English) is too strong for one-to
Analysis
TL;DR
- The study challenges three foundational assumptions in Voynich manuscript analysis: that glyphs function as letters, that token boundaries correspond to words, and that all blanks serve as word spaces
- Order in Voynichese is concentrated at token edges and graded boundaries rather than in token succession, with edge glyphs sharing 0.2 bits of mutual information—higher than any prose control
- Glyph regularity (conditional entropy 2.7 bits vs. ~3.5 for Latin/Italian/English) is too strong for one-to-one substitution ciphers and instead resolves onto quire-stable multi-symbol units
- Tokens exhibit extremely weak predictive order (under 1% of token entropy vs. 2-10% in controls), while 70% of types are hapaxes—far exceeding control corpora (41-60%)
- Published Voynich-imitating ciphers and self-citation generators reproduce low entropy and weak token order but fail to replicate edge-glyph coupling or the open vocabulary profile, suggesting the manuscript's structure is distinct from known fabrication methods
Why It Matters
This research applies rigorous computational linguistics and information-theoretic methods to one of history's most famous undeciphered texts, demonstrating that conventional analytical frameworks may be fundamentally misaligned with the manuscript's actual structure. For AI and NLP researchers, it highlights the importance of validating structural assumptions before applying standard linguistic models to unknown or artificial writing systems.
Technical Details
- Methodology: Tested three assumptions against the Zandbergen-Landini transliteration using matched prose, cipher, and pseudo-text controls with quire-level resampling to account for manuscript structure
- Entropy analysis: Conditional entropy of glyphs measured at 2.7 bits versus ~3.5 bits for Latin, Italian, and English controls, indicating higher regularity than natural languages and ruling out simple substitution ciphers
- Token order measurement: Predictive information from one token to the next falls below 1% of token entropy, significantly lower than all matched controls (2-10%), while edge glyphs show 0.2 bits mutual information exceeding any prose control
- Blank/space analysis: Blanks separate into two regimes; uncertain separators behave like word-internal junctures, are physically narrower (AUC 0.905 from image coordinates), and are crossed by learned units even when all spaces are erased before learning
- Control comparisons: A published Voynich-imitating cipher and a self-citation text generator reproduced low entropy, unit scale, weak token order, and null substitution attack results, but neither replicated edge-glyph coupling or the hapax-rich open vocabulary (70% singletons vs. 41% and 59-60% in controls)
Industry Insight
- Researchers working with unknown or artificial scripts should treat structural assumptions (letter/token/space mappings) as hypotheses requiring empirical validation rather than defaults, especially when applying NLP techniques to non-standard text
- The failure of both cipher-based and generative fake-text controls to fully replicate the manuscript's profile suggests that detecting sophisticated forgeries requires multi-dimensional analysis beyond entropy and n-gram statistics alone
- The quire-stable multi-symbol unit scale identified in this work points to a hierarchical organization that could inform how AI systems approach segmentation and unit discovery in low-resource or artificially constructed languages
Disclaimer: The above content is generated by AI and is for reference only.