Abhishek S.
Shipping in public. Listening in private.

Abhishek

I lead women’s Indo-Western & Premium at Max Fashion. I also wrote the AI that runs the buying floor.

Rare profile. Category operator who ships production code.

Senior Buying Leader · Max Fashion Women’s Indo-Western & Premium · 530+ India stores NIFT ’12 · Twelve years on the floor

abhishek@bengaluru ~ %
>role: senior buying lead
>dept: women’s indo-western + premium
>floor: 530+ stores india

William Friedman and the Voynich Manuscript — The Greatest Cryptanalyst's Greatest Failure

William F. Friedman (1891–1969) is arguably the most accomplished cryptanalyst in American history. He led the team that broke Japan's Purple cipher — the most complex cipher machine the Imperial Japanese Army had deployed — before Pearl Harbor, providing the Allies with "MAGIC" intelligence throughout the Second World War. He essentially founded the discipline of American signals intelligence, wrote the textbooks, trained the personnel, and coined the term cryptanalysis itself. He invented the index of coincidence — the statistical test that became the foundational diagnostic tool of 20th-century cryptanalysis.

He spent four decades trying to crack the Voynich Manuscript. He never got close.

That failure is one of the most instructive episodes in the history of knowledge. It reveals, with unusual precision, the difference between pattern recognition (what cryptanalysis does well) and pattern validation (what requires external ground truth) — a distinction that now haunts AI systems in exactly the same way.

Who was Friedman?

Friedman came to cryptanalysis via genetics, not mathematics. He was hired in 1915 by industrialist George Fabyan to work on Francis Bacon's claimed cipher hidden in Shakespeare's texts — a project Fabyan was financing as a private crusade. Friedman applied statistical methods to the question, concluded there was no cipher, and in doing so invented methods that became the foundation of modern cryptanalysis. He spent the rest of his professional life at the U.S. Army Signal Corps and its successors, culminating in a founding role at the NSA.

His signal achievement was the Purple analysis. The Purple machine was an electromechanical cipher device that used telephone stepping switches rather than the Enigma's rotors. Friedman's team reconstructed its complete wiring diagram purely through statistical traffic analysis — without ever seeing the machine itself. When the U.S. captured a Purple machine after Pearl Harbor, every detail matched.

He was also chronically haunted by the Voynich Manuscript, which he had encountered through bibliophile channels in the 1920s. Colleagues noted he sometimes seemed to believe cracking it was the true test of his abilities.

The First Study Group (1944–1946)

In 1944, with the war nearing its end, Friedman organized what became known as the First Voynich Manuscript Study Group (FSG) — an unofficial after-hours club of U.S. Army cryptanalysts meeting at Arlington Hall, Virginia. Members included Prescott Currier (who would later make the most significant structural discovery), John Tiltman (on secondment from British signals intelligence, and perhaps the second most accomplished cryptanalyst of his era), and Elizebeth Smith Friedman, William's wife, herself a formidable independent cryptanalyst.

The FSG's most significant technical innovation was also one of the earliest applications of computing to humanistic scholarship: they transcribed every line of the Voynich Manuscript to IBM punch cards. This made the text machine-readable — the first time a scholarly artifact had been prepared for computational analysis in this way. The IBM card coding sheets divided the transcription into 30-character blocks, each with a 3-digit serial number (001–516). For the mid-1940s, this was genuinely avant-garde.

The transcription required a working alphabet. Friedman's group developed what became the EVA (European Voynich Alphabet) precursor — a set of Latin letter equivalents for Voynich glyphs, creating the mapping infrastructure that every subsequent researcher has depended on.

What Friedman actually tried

The punch card transcription enabled Friedman and his colleagues to run the standard diagnostic battery:

1. Index of Coincidence (IC). Friedman invented the IC in 1920 as a test for cipher type. A random letter sequence has an IC of ~0.038; English has ~0.065; French has ~0.074. A monoalphabetic substitution cipher has the same IC as the underlying language (because letter frequency distributions survive substitution). A Vigenère polyalphabetic cipher has IC near random. Voynichese has an IC of roughly 0.063–0.068 — in the range of natural language, ruling out simple random padding, but also ruling out the classic Vigenère attack.

2. Frequency analysis. Standard letter-frequency profiles. Voynich glyphs show decidedly non-uniform frequency distributions, consistent with language — but the specific distributions don't match any known European language of the period, even under attempted letter-mapping.

3. Bigram and trigram analysis. Contact tables showing which glyph pairs and triples appear, how often, and in what positions. This is how Tiltman later identified the Currier A / Currier B distinction — two statistically distinct text populations, concentrated in different manuscript sections, written by different hands. This finding (which emerged from the FSG data even if formally published later) became the hardest constraint any future theory had to explain.

4. Pattern word and repeated phrase identification. The group systematically catalogued repeated word sequences, looking for anything that might serve as a "crib" — a known-plaintext fragment whose cipher equivalent would anchor decryption. They found extensive repetition (Voynich words repeat more than typical natural languages would), but couldn't anchor any of it to known text.

5. Index of entropy and redundancy. By the early 1950s, Shannon's information theory provided new tools. Entropy measures and redundancy calculations confirmed what the IC had suggested: Voynichese has language-like entropy at the word level but anomalously low entropy at word boundaries — a pattern that didn't match any known cipher type in Friedman's experience.

None of these tests yielded a crack. Every technique designed for known cipher types returned anomalous results that pointed toward an unusual text but did not point toward a solution.

The key reason for failure: no crib

Every successful cipher break in history has depended on known plaintext — a crib that anchors statistical analysis to confirmed meaning. Enigma was broken because German operators used predictable phrases ("Heil Hitler" at message end, weather reports at fixed formats, "nothing to report"). Purple was broken partly because Japanese diplomatic messages followed protocol templates.

Voynichese has zero verified vocabulary. Without a single confirmed word, statistical methods can identify anomalies and rule out known cipher types, but cannot converge on a solution — because any pattern consistent with the observed statistics can equally well support dozens of incompatible interpretations. Friedman's tests produced a fingerprint of the text; they couldn't identify whose fingerprint it was.

This is not a failure of technique. It's an epistemological limit. The problem is not solvable by more or cleverer cryptanalysis if the crib-gap remains.

Second Study Group (1962–1963)

Friedman organized a Second Study Group in 1962, with access to early computing resources and the benefit of post-Shannon information theory. John Tiltman continued working independently throughout this period. The second group reached no new conclusions beyond what the first had established.

Tiltman published his own analysis in 1968, one year before Friedman died. He noted the Voynich showed "word-like" segmentation consistent with a real language — but no statistical technique he knew could distinguish a natural language from a well-constructed synthetic one at this level of analysis.

Friedman's sealed conclusion

Rather than admit defeat publicly, Friedman composed his opinion as an anagram embedded in an article about Chaucer's Canterbury Tales. The anagram read: "I put no trust in anagrammatic acrostic cyphers, for they are of little real value—a waste—and may prove nothing."

After his death in 1969, the editor of Philological Quarterly opened the sealed envelope bearing the anagram's solution:

"The Voynich Manuscript was an early attempt to construct an artificial or universal language of the A-Priori type."

Friedman's conclusion was that the manuscript was not a cipher at all — not a secret message encoding something that could be decoded — but rather an invented language, consciously constructed according to formal rules, of the kind Francis Bacon and John Wilkins had proposed and attempted in the 17th century (notably An Essay Towards a Real Character and a Philosophical Language, 1668). In Friedman's reading, there was no plaintext to recover, because the language was the text.

Why his conclusion is now minority opinion

The a priori language hypothesis was dominant for several decades after Friedman's death. It is now a minority position, for three reasons:

First, the 2026 NLP analysis presented at the First International Conference on the Voynich Manuscript (University of Malta) found genuine prefix-root-suffix word architecture in Voynichese — a morphological structure consistent with natural inflectional languages, but inconsistent with the kind of rigid logical categories A-priori languages like Wilkins' actually have. A-priori languages tend toward simpler, more mechanical structure than Voynichese displays.

Second, the Naibbe cipher hypothesis (Greshko, Cryptologia, November 2025) demonstrated that a verbose homophonic substitution cipher using playing cards and dice could reproduce all Voynichese statistics — including the low h2 entropy at word boundaries — without positing an artificial language. The cipher hypothesis is now mechanically viable.

Third, the five-scribe finding (Fagin Davis, 2024) — that the manuscript was produced by five distinct hands — argues against a single lone eccentric constructing a language, and toward institutional production of a secret document. Secret documents tend to use ciphers, not artificial languages.

Friedman's conclusion may still be correct. But the field has moved on.

What Friedman's failure teaches

The index of coincidence test, frequency analysis, and bigram statistics are extraordinarily powerful on known cipher types. They are nearly powerless on unknown ones. The test tells you what something is not; it rarely tells you what something is.

More precisely: what Friedman's toolbox can do is eliminate hypotheses. It eliminated monoalphabetic and standard polyalphabetic ciphers immediately. It eliminated the possibility of random text. It identified the Currier A/B split. This is valuable negative knowledge — but negative knowledge does not accumulate into a decipherment.

The Voynich is the epistemological limit case: a text where the available evidence constrains the space of possible interpretations but does not uniquely select one. AI systems in 2025–2026 face the same wall when attempting to decode it — not because the models are weak, but because more pattern recognition cannot substitute for the external anchor that is missing. See concept voynich theories for why GPT-class models and neural MT have also failed, and by the same logic.

Key Facts

See Also