The Enigma of the Voynich Manuscript: A Journey Through Time and Technology

The Enigma of the Voynich Manuscript: A Journey Through Time and Technology

The Voynich Manuscript, an ancient text filled with mysterious illustrations and an unreadable script, has puzzled scholars and enthusiasts for centuries.

Most people assume that if artificial intelligence (AI) ever cracked the code of this enigmatic manuscript, it would be heralded as a monumental breakthrough in the world of linguistics and cryptography.

However, the reality is far more complex and intriguing than any headline might suggest.

In 2017, a team of researchers announced that their algorithm had found Hebrew hidden within the manuscript’s 600 years of unreadable text.

This claim sent shockwaves through the academic community and captured the attention of the media.

Yet, as medieval scholars closely examined the findings, they uncovered a much different story, revealing the limitations and challenges inherent in deciphering this cryptic work.

A Brief History of the Voynich Manuscript

The story of the Voynich Manuscript begins in 1912 when rare book dealer Wilfred Voynich discovered it among old chests of manuscripts at a Jesuit property outside Rome.

Voynich described the book as an “ugly duckling” filled with unreadable writing and illustrations of plants that no one could identify.

However, its history predates Voynich’s discovery by several centuries.

A letter tucked inside the manuscript, written by Prague scholar Johannes Marcus Marcy in 1666, claimed the manuscript once belonged to Emperor Rudolph II.

Marcy sent the book to Jesuit scholar Athanasius Kircher, hoping he could decode its contents.

Unfortunately, Kircher’s efforts were in vain, and the manuscript vanished from record for over two hundred years until Voynich brought it to light.

After failing to sell it during his lifetime, the manuscript eventually found a permanent home at Yale’s Beinecke Rare Book and Manuscript Library in 1969, where it remains today under the catalog number MS408.

Picture background

The Manuscript Nobody Can Read

The Voynich Manuscript consists of roughly 240 pages, originally believed to be closer to 272.

These pages are filled with a flowing script that resembles no known alphabet on Earth.

Accompanying the text are strange illustrations of plants that do not match any known species, star charts that do not correspond to any recognized celestial patterns, and human figures whose significance remains unclear.

Paleographers, specialists who study ancient handwriting, believe the text was produced by multiple scribes who utilized an invented alphabet, as evidenced by the distinct letter habits observed throughout the manuscript.

For centuries, this combination of factors has attracted cryptographers, chemists, and botanists alike, yet none have managed to produce a confirmed translation.

The Rise of Modern Computing

The landscape began to shift with the advent of modern computing and natural language processing.

In February 2014, language professor Steven Bax claimed to have identified around ten individual words in the manuscript by matching them to the illustrated plants and stars adjacent to the text.

Bax’s method echoed the approach used to crack ancient Egyptian hieroglyphs, focusing on proper names and building outward from there.

His proposed readings included plant names like juniper and coriander, as well as a possible reference to the constellation Taurus.

Despite the dramatic media coverage heralding his findings as a breakthrough, Bax himself remained cautious, labeling his discoveries as provisional rather than definitive.

Other researchers, including Gordon Rug, who would later publish his own competing theory, scrutinized Bax’s work and found it unconvincing, noting that his method was overly broad and could potentially fit any known or unknown language family.

As Bax continued to update his findings, the excitement surrounding his claims gradually faded without gaining acceptance among the broader linguistic community.

Picture background

The AI Breakthrough That Wasn’t

Fast forward to 2017, when a team led by computer scientist Gregor Condrach from the University of Alberta took a bold step forward in the quest to decipher the Voynich Manuscript.

Building on an earlier idea proposed by Gordon Rug, the researchers hypothesized that the manuscript’s unusual word patterns might resemble “alphags,” a concept where letters are rearranged into alphabetical order.

Using modern natural language processing tools, Condrach and his graduate student Bradley Hower trained their algorithm on the Universal Declaration of Human Rights, translated into 380 different languages.

They tested whether the system could accurately identify the source language of scrambled versions of the text.

To their surprise, the algorithm achieved an impressive success rate of approximately 97%, even when the letters within each word were shuffled.

With this newfound capability, the researchers directed their algorithm at the actual text of the Voynich Manuscript, initially suspecting that the text might be Arabic.

However, the algorithm flagged Hebrew as the most likely candidate, generating a wave of headlines proclaiming the mystery solved.

Experts Respond: Not So Fast

Within days of the announcement, experts began to voice their skepticism.

Condrach himself clarified that his team never claimed to have definitively deciphered the manuscript; rather, they viewed their results as an early signal worth investigating.

The real challenge arose when the researchers attempted to translate the algorithm’s findings into coherent sentences.

According to their own admission, the first translated line was not coherent and required manual spelling corrections and a pass through Google Translate to read like English.

Lisa Fagan Davis, executive director of the Medieval Academy of America, expressed her concerns, pointing out that an algorithm trained on modern languages cannot reliably identify the language of a document from the 1400s.

The differences in grammar, spelling, and vocabulary would have been significant, especially in a technical text like the Voynich Manuscript.

 

Medievalist Damian Fleming echoed these sentiments publicly, emphasizing the limitations of the algorithm’s approach.

Condrach later acknowledged that the initial reactions from Voynich specialists were lukewarm at best, but this setback did not halt computational research into the manuscript.

Instead, it prompted researchers to refine their focus, shifting towards testing whether the text behaves like a language rather than attempting to identify which one it might be.

A Different Approach: Hidden Markov Models

Years before Condrach and Hower’s efforts, computational linguist Kevin Knight and researcher Shravana Reddi took a more cautious approach back in 2011.

Rather than attempting to name a language, they built a hidden Markov model, a statistical tool designed to uncover hidden patterns within visible data.

The model was applied to the entire transcribed manuscript available at the time, which consisted of approximately 225 pages and 8,100 words.

The goal was to categorize each page according to the two writing styles identified by cryptologist Prescott Courier in the 1970s, known as Courier Language A and B.

The model successfully calculated the degree to which each folio leaned toward one style or the other, transforming a decades-old hypothesis into a precise statistical map.

This map aligned closely with the manuscript’s illustrated sections, revealing that Language A dominated the herbal pages, while Language B prevailed in the biological and recipe sections.

Picture background

The Statistical Debate

Amidst these findings, researchers stumbled upon an intriguing detail: the letters within Voynich words exhibited an unusually strict predictable pattern compared to real languages like English or Latin, a phenomenon known as low entropy.

This observation divided researchers into two camps for the next decade.

One camp interpreted it as evidence of a genuine but unusual language, while the other viewed it as a warning sign of something artificial.

Determining which perspective held more weight required testing the manuscript against a more extensive dataset than any single study could provide.

In June 2013, physicist Marcelo Montemurro led a team that published findings in the journal PLOS ONE, treating the entire manuscript as a vast network of interconnected words.

Using tools from information theory, which measures the amount of information contained within any message, the team sought to determine whether the manuscript exhibited the characteristics of a real human language.

Real languages tend to cluster meaningful words—nouns and verbs—around specific topics while smaller connector words, such as “and” or “the,” scatter evenly throughout the text.

When Montemurro’s team applied this test to the Voynich manuscript, certain words behaved similarly to meaningful content words in genuine languages, clustering closely around specific sections rather than appearing randomly across the pages.

Such structured behavior is not typically produced by random gibberish.

Picture background

Competing Studies and Findings

In 2013, another team led by Diego Amancio at the University of Sao Paulo published a study comparing the manuscript directly against a large library of real texts written in dozens of confirmed human languages.

Their approach involved measuring three separate categories of statistical behavior side by side.

First, they assessed whether the manuscript adhered to Zipf’s law, a mathematical principle observed across nearly all known human languages, where a small number of words appear frequently while the majority are used rarely.

Second, they created complex network models, treating each word as a connected point, similar to Montemurro’s method, to measure the overall shape of that network against real language networks.

Lastly, they analyzed the text as a time series, tracking how words repeat or cluster together as a reader progresses through the pages, akin to techniques used to study stock market patterns.

Across all three tests, the manuscript consistently fell within the range produced by real human languages rather than the vastly different range generated by random noise.

However, the researchers were careful to frame their conclusions with caution, acknowledging that their statistical framework could not identify the specific language of the text or prove that it contained genuine hidden meaning.

What they could assert with confidence was that the Voynich Manuscript exhibited statistical compatibility with being a genuine natural language text rather than meaningless filler.

The Evolution of Research

The convergence of findings from two independent research teams using different mathematical tools bolstered the argument for the manuscript’s linguistic significance.

This growing consensus prompted a third group to undertake an even larger comparison, measuring the manuscript against a vast array of the world’s written languages.

In 2020, Yale linguists Luke Lindaman and Clare Bowen expanded the comparison problem significantly.

Picture background

Instead of analyzing the manuscript against a limited selection of texts, they constructed two new digital collections: one featuring cleaned text from Wikipedia articles in various modern languages and another encompassing a historical collection spanning centuries and language families.

Using a measure known as conditional character entropy, which tracks how predictable each letter is once the preceding letter is known, they confirmed and refined the observations made by Ready and Knight regarding the manuscript’s unusually low entropy.

Lindaman and Bowen did not treat this finding as definitive proof of a hoax, but rather as evidence that certain historical encoding methods, such as systems combining two letters into one written symbol (known as bigraphs), could produce a genuine language’s low entropy score.

Their work has become a widely cited reference in the field, providing fresh ammunition for researchers who argue that the manuscript may not have needed a real language behind it at all.

A Discovery in the Margins

As the statistical arguments continued to unfold, a separate and more tangible discovery emerged in 2024.

Lisa Fagan Davis revisited multispectral images of the manuscript, originally captured a decade earlier by the Lazarus Project, and examined the faint marks found in the margins of certain pages.

These marks, too faded to be discerned by the naked eye, revealed a string of ordinary Roman letters likely left by someone attempting to decipher the manuscript long before modern computational tools existed.

While this discovery did not translate any words from the actual Voynich script, it indicated that at least one earlier reader had engaged seriously with the text as a puzzle worth solving.

This find carries significance beyond mere curiosity, as it confirms that the manuscript has been treated as a genuine enigma for centuries, predating AI and statistical modeling.

The Ongoing Debate: Language or Nonsense?

Despite the compelling evidence supporting the idea that the Voynich Manuscript may contain meaningful language, not all computational researchers agree.

Psychologist Gordon Rug, working at Keele University, proposed a theory suggesting that the manuscript could have been produced using a low-tech method available in the 15th century.

According to Rug’s hypothesis, a device called a cardan grill, combined with tables of pre-arranged syllables, could generate an endless stream of new invented words without requiring any knowledge of a real language or encoding actual hidden meaning.

Picture background

Rug constructed a working demonstration of this method, arguing that it could reproduce several of the statistical characteristics researchers had identified in the manuscript, including low entropy patterns.

In 2016, Rug published a follow-up study testing whether this table-based method could replicate even more statistical features, yielding promising results.

A separate challenge came from Torsten Tim and physicist Andreas Scher, who published a detailed study proposing their own mechanical process for generating the manuscript’s text.

By closely analyzing the patterns of similarly spelled words across the pages, they claimed to have identified the likely methods used to mechanically produce new words throughout the text.

If their theory holds, it would imply that the manuscript’s text was never intended to encode real meaning, and the statistical signals detected by researchers could be side effects of a mechanical generation process rather than evidence of a genuine language.

The Unresolved Divide

Neither the hoax theory nor the mechanical generation theory has fully convinced the wider research community.

Proponents of the natural language interpretation argue that a purely mechanical process struggles to account for some of the complex semantic clustering patterns documented by Montemurro’s team, as well as the correlations between specific words and the illustrations on adjacent pages.

This unresolved divide among researchers keeps attracting new, often less rigorous attempts to claim a final solution to the manuscript’s mysteries.

The cycle of bold claims followed by swift retractions continues to play out in the academic landscape.

In 2019, the University of Bristol issued a press release announcing that researcher Gerard Chesher had identified the manuscript’s text as an extinct language he dubbed proto-Romance.

Almost immediately, multiple specialists, including Lisa Fagan Davis, publicly criticized the claim as unreliable, leading Bristol to retract its press release within a day.

Similar patterns have emerged in less formal venues, with several self-published papers on academic sharing sites claiming to have solved the manuscript using AI-assisted methods.

These include claims that the text is actually old Turkish or describes medical treatments related to astrology, supposedly authored by a specific Swiss German monk who traveled extensively through Asia and Africa.

None of these claims have appeared in peer-reviewed journals or received independent verification from researchers who have dedicated decades to studying the manuscript.

They all exhibit the same warning signs that led to the downfall of Chesher’s proto-Romance theory in 2019.

The Importance of Rigorous Research

In contrast, the genuine computational research surrounding the Voynich Manuscript stands apart from these sensational claims.

Ready and Knight’s hidden Markov model, Montemurro’s information-theoretic network analysis, Amancio’s statistical physics comparisons, and Lindaman and Bowen’s large-scale entropy collections were all published in peer-reviewed venues.

Each study clearly articulated the specific limitations of its findings and refrained from claiming to have translated any actual sentences of meaning.

Even Condrach, after his AI results gained media attention, was careful to clarify that he never claimed a full decipherment of the manuscript.

This careful and measured approach to research is far less exciting than headlines proclaiming the mystery solved, but it is the only version of this research that has withstood scrutiny from experts over the years.

Conclusion: The Mystery Remains

As it stands, the Voynich Manuscript is a genuine radiocarbon-dated document from the early 1400s, with a strange history that traces its ownership from Emperor Rudolph II’s court in Prague to a Jesuit library outside Rome, ultimately landing at Yale’s Beinecke Rare Book and Manuscript Library in 1969.

It has resisted every attempt at translation for over a century, from Steven Bax’s provisional plant name readings in 2014 to the more audacious AI-driven Hebrew claim in 2017.

While the AI attempt by Condrach and Hower pointed towards Hebrew hidden within an alphagram cipher, their own translated output required manual corrections and Google Translate to achieve coherence, leading medieval scholars like Lisa Fagan Davis to reject the method outright.

Condrach himself later acknowledged that his team never claimed to have achieved a full decipherment.

Simultaneously, a careful line of statistical research, spanning from Ready and Knight in 2011 to Montemurro and Amancio’s teams in 2013 and Lindaman and Bowen in 2020, has consistently identified patterns indicative of real human language.

At the same time, researchers such as Gordon Rug, Torsten Tim, and Andreas Scher have presented working demonstrations showing that simple mechanical methods could replicate several of the same statistical features without any underlying real language.

The discovery of faded marginal writing in 2024 further underscores that this manuscript has been regarded as a genuine puzzle for centuries, long before modern tools emerged.

The truth behind the Voynich Manuscript remains elusive and complex, straddling the line between the possibility of a hidden language and the potential for an elaborate imitation.

As of now, no computer, algorithm, or AI model has produced a confirmed translation of a single sentence from this manuscript, leaving the debate open and unresolved.

The gap between genuine published research and sensational claims of victory highlights the importance of rigorous academic inquiry in the face of enticing headlines.

While the enigma of the Voynich Manuscript endures, it continues to captivate the imagination of scholars and enthusiasts alike, ensuring that its mysteries will be explored for years to come.

Disclaimer: This content may be created by Al for entertainment purposes. Any resemblance to real persons, events, or places is coincidental.

Disclaimer: This story is fictional and created for entertainment purposes only. Any names, characters, places, or events are fictitious or used fictitiously. No real person or organization is intended to be portrayed.

Recommended for You

View Archive arrow_forward