← Laboratory

EXP-0005 — Exact repetition vs word-shuffled surrogates (first null-model test)

completed Deterministic computation

Question

Does the corpus contain more exact repeated word sequences than surrogates that preserve its exact word inventory and ayah lengths but destroy word order?

Hypothesis

One-sided: the observed count of distinct repeated 8-token sequences exceeds the word-shuffle null. Declared PRIMARY endpoint: distinct repeated 8-token sequences (letters-basic-v1). Declared SECONDARY endpoint: distinct repeated 5-token sequences. Both endpoints, the surrogate class, replicate count, and base seed are fixed in this spec before execution (D-044).

Counting policy

letters-basic-v1; whitespace tokens; surrogate class word-shuffle-v1 preserves the exact global word multiset and per-ayah token counts. Interpretation scope is limited to unusualness under this surrogate class; word-shuffle nulls only establish ordered/formulaic composition, a property of essentially all authored texts.

Results — permutation test vs declared surrogate class

Surrogate: word-shuffle-v1 (preserves the exact word inventory and per-ayah lengths; destroys all word order).

Primary endpoint (declared in spec) — distinct repeated 8-token sequences

observed 611 · null mean 0 ± 0 · null range [0, 0] over 1,000 surrogates · one-sided (greater) · p = 0.000999 (add-one estimator; 1/1001 is the floor) · base seed 20260731 · full null sample persisted in the result file

Secondary endpoint — distinct repeated 5-token sequences

observed 2,302 · null mean 0.001 ± 0.03162 · null range [0, 1] over 1,000 surrogates · one-sided (greater) · p = 0.000999 (add-one estimator; 1/1001 is the floor) · base seed 20260731 · full null sample persisted in the result file

Inference scope

The observed corpus contains far more exact repetition than word-shuffled surrogates. This quantifies that the text is ordered, formulaic composition rather than a bag of words — a property shared by essentially all authored texts, and the same test would separate any coherent book from its shuffle. Its value is as a calibrated baseline for future comparisons against matched real texts (prose, poetry — §13.5), not as evidence for any claim about authorship.

Reproducibility record

Corpus
quran:hafs-kufan:v1
Method version
repetition-vs-null-v1
Results sha256
812b2afca8d12ca4713c7c675b66af74…
Environment
{"platform":"Linux-6.18.5-x86_64-with-glibc2.39","python":"3.11.15"}
Ran
2026-07-31T03:01:25Z → 2026-07-31T03:03:08Z

Re-run with python3 services/research-worker/scripts/run_experiment.py data/research/specs/EXP-0005.json. Input checksums are recorded in the result file.