EXP-0005 — Exact repetition vs word-shuffled surrogates (first null-model test)
completed Deterministic computation
Question
Does the corpus contain more exact repeated word sequences than surrogates that preserve its exact word inventory and ayah lengths but destroy word order?
Hypothesis
One-sided: the observed count of distinct repeated 8-token sequences exceeds the word-shuffle null. Declared PRIMARY endpoint: distinct repeated 8-token sequences (letters-basic-v1). Declared SECONDARY endpoint: distinct repeated 5-token sequences. Both endpoints, the surrogate class, replicate count, and base seed are fixed in this spec before execution (D-044).
Counting policy
letters-basic-v1; whitespace tokens; surrogate class word-shuffle-v1 preserves the exact global word multiset and per-ayah token counts. Interpretation scope is limited to unusualness under this surrogate class; word-shuffle nulls only establish ordered/formulaic composition, a property of essentially all authored texts.
Results — permutation test vs declared surrogate class
Surrogate: word-shuffle-v1 (preserves the exact word inventory and per-ayah lengths; destroys all word order).
Primary endpoint (declared in spec) — distinct repeated 8-token sequences
observed 611 · null mean 0 ± 0 · null range [0, 0] over 1,000 surrogates · one-sided (greater) · p = 0.000999 (add-one estimator; 1/1001 is the floor) · base seed 20260731 · full null sample persisted in the result file
Secondary endpoint — distinct repeated 5-token sequences
observed 2,302 · null mean 0.001 ± 0.03162 · null range [0, 1] over 1,000 surrogates · one-sided (greater) · p = 0.000999 (add-one estimator; 1/1001 is the floor) · base seed 20260731 · full null sample persisted in the result file
Inference scope
The observed corpus contains far more exact repetition than word-shuffled surrogates. This quantifies that the text is ordered, formulaic composition rather than a bag of words — a property shared by essentially all authored texts, and the same test would separate any coherent book from its shuffle. Its value is as a calibrated baseline for future comparisons against matched real texts (prose, poetry — §13.5), not as evidence for any claim about authorship.
Reproducibility record
- Corpus
- quran:hafs-kufan:v1
- Method version
- repetition-vs-null-v1
- Results sha256
- 812b2afca8d12ca4713c7c675b66af74…
- Environment
- {"platform":"Linux-6.18.5-x86_64-with-glibc2.39","python":"3.11.15"}
- Ran
- 2026-07-31T03:01:25Z → 2026-07-31T03:03:08Z
Re-run with python3 services/research-worker/scripts/run_experiment.py data/research/specs/EXP-0005.json. Input checksums are recorded in the result file.
