← Laboratory

EXP-0004 — Root rank-frequency distribution

completed Deterministic computation

Question

What does the rank-frequency distribution of triliteral roots look like across the corpus?

Hypothesis

Descriptive baseline; no directional hypothesis. Provides the frequency table behind the rare-root graph threshold (D-028) and future lexical-trajectory work (Program C).

Counting policy

Root tags from morphology dataset quran-morphology-8f38b39; words without a root tag (particles, some pronouns) are excluded by construction.

Results — root rank-frequency

1651 distinct roots over 50,268 root-tagged words. Log-log slope -1.3958 (Ordinary least squares on log(frequency) vs log(rank), roots with frequency >= 2. Descriptive statistic only; no inferential claim.)

RankRootWordsAyahs
1أله2,8511,879
2قول1,7221,383
3كون1,3901,176
4ربب980871
5أمن879723
6علم854728
7قوم660597
8أيي597562
9أتي549486
10كفر525465
11بين523454
12شيأ519449
13رسل513429
14يوم475437
15أرض461440
16سمو381352
17كلل377355
18عذب373336
19عمل360313
20جعل346311
21رحم339313
22أنس338317
23رأي328297
24كتب319279
25هدي316268
26ظلم315290
27نفس298270
28قبل294282
29نزل293257
30ذكر292264

Reproducibility record

Corpus
quran:hafs-kufan:v1
Method version
root-frequency-v1
Results sha256
bf03c335105f31ad2f4fb83c48d512b1…
Environment
{"platform":"Linux-6.18.5-x86_64-with-glibc2.39","python":"3.11.15"}
Ran
2026-07-31T02:31:56Z → 2026-07-31T02:31:56Z

Re-run with python3 services/research-worker/scripts/run_experiment.py data/research/specs/EXP-0004.json. Input checksums are recorded in the result file.