A Statistical Turing Test for the Voynich Manuscript: Evaluating Eight Quantitative Linguistic Properties Across Nine Generative Model Families

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

We introduce a Statistical Turing Test (STT) protocol for the Voynich manuscript and use it to evaluate whether modern generative models can produce text that is statistically indistinguishable from the manuscript across eight simultaneously applied quantitative metrics: Zipf’s law slope, character-level conditional entropy, word-level conditional entropy, standardised type-token ratio, word-length distri bution, normalised pointwise mutual information of bigrams, positional vocabu lary divergence, and Currier dialect replicability. Eighteen model variants spanning nine families — uniform sampling, unigram, Cardan grille simulation, character level Markov chains (orders 1–5), word-level Markov chains (orders 1–3), a two state hidden Markov model, a character-level LSTM, GPT-2 fine-tuned on the manuscript, and a section-stratified bigram sampler — are evaluated over 30 in dependent runs each. No model passes all eight criteria simultaneously; the best model (section-stratified interpolated bigram, L8b) passes 7/8. The Cardan grille simulation proposed by Rugg (2004) passes only 1/8, directly quantifying its sta tistical inadequacy. We find that positional vocabulary divergence (M7) is the met ric most resistant to stationary generators: all models that draw from a single vo cabulary distribution fail this test. Zipf’s law slope is the single hardest property to replicate: every model tested produces a steeper rank-frequency slope than the manuscript, suggesting an anti-concentration mechanism in the Voynich text that goes beyond frequency-weighted sampling. A complexity analysis shows that increas ing generative model complexity does not monotonically improve statistical fidelity; the minimum-complexity threshold for a full STT pass has not been reached in the tested range. We release all code and generated corpora to support reproducibility.

Article activity feed