← Back to list

The Voynich Manuscript — The Book Nobody Has Read

In 1912, a Polish-American antiquarian book dealer named Wilfrid Voynich was browsing a collection of manuscripts at the Villa Mondragone…

J Gray in What If AI Investigated…? · 2026-07-02 07:47 · 0 claps · 9.1 min read paywalled
#voynich-manuscript #what-if-ai-investigated #cryptography #medieval #linguistics
Open on Medium ↗
Wiki topics: RAG · RAG & Retrieval CRY · Crypto & Web3 HIS · History LNG · Linguistics & Language 🔒 · Cybersecurity

The Voynich Manuscript — The Book Nobody Has Read

Two hundred and forty pages of consistent, structured, illustrated text. In a script nobody has read. In a language nobody has identified. Written on vellum carbon-dated to 1404–1438. All images generated by author unless stated

Two hundred and forty pages of consistent, structured, illustrated text. In a script nobody has read. In a language nobody has identified. Written on vellum carbon-dated to 1404–1438. All images generated by author unless stated

In 1912, a Polish-American antiquarian book dealer named Wilfrid Voynich was browsing a collection of manuscripts at the Villa Mondragone near Frascati, Italy, when he came across something he had never seen before.

It was a codex of approximately 240 pages, written on vellum, illustrated throughout with botanical drawings, astronomical diagrams, naked figures bathing in pools connected by elaborate pipe systems and what appeared to be zodiac charts. The text accompanying these illustrations was written in a flowing, apparently consistent script that Voynich had never encountered in any other manuscript. He could not read a word of it. Nobody he showed it to could read a word of it either.

He bought it. Named it after himself. And began what would become one of the longest running and least resolved puzzles in the history of cryptography, linguistics and medieval studies.

The Voynich Manuscript is now housed at Yale University’s Beinecke Rare Book and Manuscript Library as MS 408. It has been examined by professional codebreakers, computational linguists, artificial intelligence researchers, medievalists, botanists, astronomers and a very large number of amateur enthusiasts who have each, in turn, believed they were about to solve it.

None of them has.

Established: what the object is

Several things about the Voynich Manuscript are known with reasonable certainty.

The date. In December 2009, a team led by physicist Greg Hodgins at the University of Arizona conducted radiocarbon analysis of the manuscript’s vellum, with results announced at a press conference and subsequently widely reported. The dating placed the vellum’s production between 1404 and 1438 with 95% confidence. These results were not published in a formal peer reviewed paper, the analysis was conducted for an Austrian television documentary, but they have been accepted across the research community as the best available dating evidence. The date rules out a number of proposed authors, including Roger Bacon, who died in 1292, and places the manuscript firmly in early 15th-century Europe.

The physical object. The manuscript currently contains 240 pages, though evidence of missing pages suggests it was originally longer. The text is written in brown ink with some red and blue used in illustrations. The illustrations are categorised by researchers into sections: botanical (drawings of plants, most not identifiable as real species), astronomical (circular diagrams including what appears to be a zodiac), balneological (naked figures in elaborate bathing systems connected by pipes), cosmological (fold-out circular diagrams), pharmaceutical (jars and containers accompanied by plant drawings) and recipes (text in paragraphs without illustrations). The categorisation is based on what the illustrations appear to show, not on any understanding of the accompanying text.

The script. The Voynich script contains approximately 25–30 distinct characters, depending on how researchers define character boundaries. The script runs left to right. There are no corrections or overwriting visible in the text, which either indicates an unusually confident scribe or a copying exercise from an already composed text. The script is internally consistent throughout.

The provenance chain. The manuscript can be traced with confidence to the court of Holy Roman Emperor Rudolf II in Prague, who reportedly purchased it in the late 16th century. A letter from 1666, written by Johannes Marcus Marci to the Jesuit scholar Athanasius Kircher, accompanied the manuscript and mentioned that Rudolf had paid 600 gold ducats for it. Marci sent it to Kircher in hopes he could decode it. Kircher, despite being one of the most learned men in Europe, apparently could not.

In 2024, multispectral imaging of the manuscript’s first page revealed previously hidden columns of letters, Marci’s own attempts to decrypt the text in the 1660s, written in lighter ink that had become invisible over centuries. He also got nowhere.

The botanical section contains drawings of approximately 130 plants. Botanists have identified none of them with certainty as real species.

The botanical section contains drawings of approximately 130 plants. Botanists have identified none of them with certainty as real species.

Established: the statistical properties of the text

The most significant body of research on the Voynich Manuscript is not decipherment attempts but statistical analysis of the text’s structure. Several robust findings have emerged from decades of computational work.

Zipf’s law compliance. In 1935, linguist George Zipf observed that in any natural language, the frequency of a word is inversely proportional to its rank: the most common word appears roughly twice as often as the second most common and three times as often as the third. This power law holds across languages as different as English, Mandarin and ancient Sumerian. When researchers apply Zipf’s law to the Voynich text, treating its distinct character sequences as words, the text complies. The distribution follows the expected power law curve.

This is significant. It means the text does not behave like random noise. But it is not diagnostic: researchers have since shown that certain simple mechanical processes can produce Zipf compliant text without encoding any meaning. Zipf compliance is a necessary but not sufficient condition for natural language.

Entropy. Information entropy measures the unpredictability of a sequence. Natural languages have characteristic entropy values, high enough to carry information efficiently and low enough to be learnable. Amancio et al. (2013) found that the Voynich text’s entropy values are intermediate between natural language and random text, closer to natural language than to random sequences. The text is neither random nor as constrained as a simple cipher.

Word structure. The Voynich text has consistent internal word structure: certain character combinations appear frequently at word beginnings, others at word endings. This is characteristic of agglutinative languages, which build words by adding prefixes and suffixes to roots. The word length distribution is narrower than most natural languages, which is unusual and has been cited as evidence against it being a conventional natural language or a simple vowel dropping abjad.

Syntactic regularities. Words appear to cluster in patterns consistent with syntax. Certain words appear primarily at the beginnings of lines, others at the ends. The clustering is consistent across the manuscript’s different sections, suggesting a unified compositional logic.

What these statistical properties collectively establish is that the Voynich text is not random. It has structure. The structure is language-like. The structure does not tell us whether that language-like structure encodes meaning.

Contested: what the text is

Four main hypotheses compete to explain what the Voynich Manuscript is. None has been accepted as definitive.

A natural language in an unknown or rare script. The most intuitive hypothesis is that the manuscript encodes a real language written in a custom alphabet. Proposed candidate languages have included Hebrew, Arabic, Latin, Italian and Nahuatl. The most recent and most discussed proposal is Old Turkish, advanced by Canadian electrical engineer Ahmet Ardıç and the ATA Research Group in papers presented at Turkish academic conferences in 2022 and 2023. Ardıç claims to have read over 1,000 Voynich words as Old Turkish phonetic content. Lisa Fagin Davis, Executive Director of the Medieval Academy of America and a leading Voynich researcher, described the work as one of the few solutions she had seen that is consistent, repeatable and results in sensical text, a notable qualified endorsement. But Davis also noted that no proposed solution has yet satisfied all five of her criteria for a credible decipherment. Critics argue that Ardıç’s method allows most Voynich characters to represent multiple sounds, which permits arbitrary translation rather than revealing a genuine linguistic system. The hypothesis has not been published in a mainstream peer-reviewed journal or independently validated by Voynich researchers outside the ATA group. Evidence tier: Contested.

A 2018 study by Kondrak and Hauer at the University of Alberta applied AI natural language processing to the manuscript and found that its statistical properties most closely resembled medieval Hebrew, with the researchers hypothesising a cipher involving alpha-grams, alphabetically ordered letter rearrangements with vowels dropped. Attempts to apply this system to the first ten pages produced mixed results. No coherent extended translation followed.

A cipher. The manuscript may encode a known language in a cryptographic system. Professional codebreakers from the US National Security Agency and GCHQ have worked on this hypothesis. None has produced a decipherment. The text’s entropy profile is not consistent with simple substitution ciphering of a known language, which would preserve the underlying language’s frequency distribution.

An elaborate hoax. The manuscript may be a sophisticated nonsense text, designed to look like a real encoded manuscript and deceive a buyer, possibly Rudolf II and his 600 gold ducats. Researcher Gordon Rugg demonstrated in 2004 that a Cardan grille, a simple mechanical device, could generate Zipf compliant nonsense text with the word structure properties of the Voynich manuscript. A 2025 academic paper formalised this, presenting a self-citation process that a medieval scribe could have executed without additional tools to reproduce the manuscript’s key statistical properties.

The hoax hypothesis is not proven. But it explains the manuscript’s most puzzling feature: why five centuries of expert analysis have failed to extract meaning from a text that has every structural appearance of containing some.

Glossolalia or idiosyncratic constructed language. A smaller number of researchers have proposed that the manuscript represents a personal invented language or something closer to glossolalia. These hypotheses predict essentially the same statistical properties as the hoax hypothesis while requiring a different psychological account of the manuscript’s creation.

The astronomical section contains circular diagrams that resemble zodiac charts. The symbols accompanying them have not been read.

The astronomical section contains circular diagrams that resemble zodiac charts. The symbols accompanying them have not been read.

Open: why decipherment keeps failing

The most analytically significant fact about the Voynich Manuscript is not that it hasn’t been deciphered. Difficult ciphers can resist decipherment for centuries. The significant fact is that every claimed decipherment has failed in the same way.

The pattern is consistent. A researcher proposes a reading. The reading produces a few coherent looking words or phrases. It fails to extend to the rest of the manuscript. Independent verification is not achieved. The proposed reading is set aside.

This pattern has repeated so many times that Voynich scholarship has a specific problem with it. The manuscript is unusual enough in its statistical properties that almost any sufficiently flexible decipherment hypothesis can be made to fit a portion of the text while failing on the remainder. The researcher sees meaning in the portion that fits and discounts the portion that doesn’t.

Lisa Fagin Davis, who gave the Old Turkish hypothesis one of the most qualified endorsements any proposed solution has received in recent years, nonetheless noted that no proposed solution has yet satisfied all five criteria she considers necessary for a credible decipherment. That formulation, the most generous recent professional endorsement still falling short of validation, is the accurate measure of where the field stands.

The deeper problem is structural. The text has word-like structures. Those word-like structures do not behave like words in any natural language that has been proposed as a candidate. The combination, structured enough to look like language, irregular enough to resist all language-based analysis, is precisely what the hoax hypothesis predicts. It is also what a very unusual real language might look like. The statistics cannot distinguish between these possibilities.

What AI can and cannot settle

The Voynich Manuscript has become a litmus test for the actual capabilities versus the marketing promises of artificial intelligence.

What AI has contributed: significantly more rigorous statistical analysis than was previously possible. The Kondrak and Hauer NLP study produced more systematic entropy and frequency measurements than earlier approaches. AI tools have eliminated a large number of candidate languages and cipher structures through negative results that are themselves informative.

What AI has not contributed: a working decipherment. No AI system applied to the Voynich Manuscript has produced a translation that independent scholars have validated as coherent.

The fundamental obstacle is not computational power. It is the absence of a known correspondence between Voynich characters and any known language or alphabet. Machine translation and code breaking both require an anchor, some point of known correspondence from which a system can be inferred. The Voynich text provides no such anchor. AI can describe the manuscript’s structure with extraordinary precision. It cannot extract meaning from a text whose meaning, if any, remains inaccessible.

The 2024 multispectral imaging finding, that Marci himself attempted and failed to decrypt the manuscript in the 1660s, is a useful calibration point. Marci had access to Athanasius Kircher, one of the most polymathically learned men in 17th century Europe, who had cracked other undeciphered scripts. Both of them failed. The failure of modern AI to do better than Marci and Kircher does not reflect a limitation of AI. It reflects the fundamental nature of the problem.

The Voynich Manuscript may be a real language not yet identified. It may be an elaborate hoax. It may be something that doesn’t fit cleanly into either category. What five centuries of analysis have established is that it will not yield to cleverness alone, human or computational. What would be required is external evidence: a bilingual key, a contemporary reference, an identified author. None of those things has been found.

Until one is, the manuscript remains exactly what Wilfrid Voynich found in 1912: a book that looks like it should be readable, and isn’t.

Sources: Radiocarbon dating, Greg Hodgins, University of Arizona, announced December 2009 (no formal peer-reviewed publication; results widely reported and accepted across the research community); Amancio et al., Probing the Statistical Properties of Unknown Texts, PLOS ONE (2013); Kondrak and Hauer, Decoding Anagrammed Texts Written in an Unknown Language and Script, Transactions of the Association of Computational Linguistics (2018); Rugg, Gordon, An Elegant Hoax? A Possible Solution to the Voynich Manuscript, Cryptologia (2004); Ardıç, Ahmet, ATA Research Group, conference proceedings, Niğde University (2022) and 1st International Turkish Culture Symposium (2023); Davis, Lisa Fagin, qualified endorsement of Old Turkish hypothesis and five-criteria statement; multispectral imaging findings (2024); Beinecke Rare Book and Manuscript Library, Yale University, MS 408.


메타데이터
post_id
1429451c83f2
slug
what-if-ai-investigated-the-voynich-manuscript-the-book-nobody-has-read-1429451c83f2
url
https://medium.com/what-if-ai-investigated/what-if-ai-investigated-the-voynich-manuscript-the-book-nobody-has-read-1429451c83f2
canonical_url
https://medium.com/what-if-ai-investigated/what-if-ai-investigated-the-voynich-manuscript-the-book-nobody-has-read-1429451c83f2
author_url
https://medium.com/@jamie_gray027
status
ok
fetched_at
2026-07-09 06:39:14