Agent morphology
3 candidates · 0 approved · 0 dismissed · 3 awaiting review
1743 lemmas occur exactly once in the corpus (hapax legomena) unreviewed deterministic candidate
{
"count": 1743,
"sample": [
{
"ayah": "1:7",
"form": "ٱلۡمَغۡضُوبِ",
"lemma": "مغضوب",
"word": 6
},
{
"ayah": "2:16",
"form": "رَبِحَت",
"lemma": "ربحت",
"word": 7
},
{
"ayah": "2:17",
"form": "ٱسۡتَوۡقَدَ",
"lemma": "استوقد",
"word": 4
},
{
"ayah": "2:19",
"form": "كَصَيِّبٖ",
"lemma": "صيب",
"word": 2
},
{
"ayah": "2:26",
"form": "بَعُوضَةٗ",
"lemma": "بعوضة",
"word": 9
},
{
"ayah": "2:30",
"form": "وَنُقَدِّسُ",
"lemma": "نقدس",
"word": 21
},
{
"ayah": "2:36",
"form": "فَأَزَلَّهُمَا",
"lemma": "أزل",
"word": 1
},
{
"ayah": "2:54",
"form": "بِٱتِّخَاذِكُمُ",
"lemma": "اتخاذ",
"word": 9
},
{
"ayah": "2:60",
"form": "فَٱنفَجَرَتۡ",
"lemma": "انفجرت",
"word": 9
},
{
"ayah": "2:61",
"form": "بَقۡلِهَا",
"lemma": "بقل",
"word": 18
},
{
"ayah": "2:61",
"form": "وَقِثَّآئِهَا",
"lemma": "قثائ",
"word": 19
},
{
"ayah": "2:61",
"form": "وَفُومِهَا",
"lemma": "فوم",
"word": 20
},
{
"ayah": "2:61",
"form": "وَعَدَسِهَا",
"lemma": "عدس",
"word": 21
},
{
"ayah": "2:61",
"form": "وَبَصَلِهَاۖ",
"lemma": "بصل",
"word": 22
},
{
"ayah": "2:68",
"form": "فَارِضٞ",
"lemma": "فارض",
"word": 15
},
{
"ayah": "2:68",
"form": "عَوَانُۢ",
"lemma": "عوان",
"word": 18
},
{
"ayah": "2:69",
"form": "صَفۡرَآءُ",
"lemma": "صفراء",
"word": 14
},
{
"ayah": "2:69",
"form": "فَاقِعٞ",
"lemma": "فاقع",
"word": 15
},
{
"ayah": "2:69",
"form": "تَسُرُّ",
"lemma": "تسر",
"word": 17
},
{
"ayah": "2:71",
"form": "شِيَةَ",
"lemma": "شية",
"word": 15
},
{
"ayah": "2:72",
"form": "فَٱدَّٰرَْٰٔتُمْ",
"lemma": "ادارأ",
"word": 4
},
{
"ayah": "2:74",
"form": "قَسۡوَةٗۚ",
"lemma": "قسوة",
"word": 11
},
{
"ayah": "2:74",
"form": "يَتَفَجَّرُ",
"lemma": "يتفجر",
"word": 16
},
{
"ayah": "2:85",
"form": "تُفَٰدُوهُمۡ",
"lemma": "تفاد",
"word": 18
},
{
"ayah": "2:93",
"form": "وَأُشۡرِبُواْ",
"lemma": "أشرب",
"word": 15
},
{
"ayah": "2:96",
"form": "أَحۡرَصَ",
"lemma": "أحرص",
"word": 2
},
{
"ayah": "2:96",
"form": "بِمُزَحۡزِحِهِۦ",
"lemma": "مزحزح",
"word": 17
},
{
"ayah": "2:98",
"form": "وَمِيكَىٰلَ",
"lemma": "ميكال",
"word": 8
},
{
"ayah": "2:102",
"form": "بِبَابِلَ",
"lemma": "بابل",
"word": 21
},
{
"ayah": "2:102",
"form": "هَٰرُوتَ",
"lemma": "هاروت",
"word": 22
},
{
"ayah": "2:102",
"form": "وَمَٰرُوتَۚ",
"lemma": "ماروت",
"word": 23
},
{
"ayah": "2:114",
"form": "خَرَابِهَآۚ",
"lemma": "خراب",
"word": 13
},
{
"ayah": "2:121",
"form": "تِلَاوَتِهِۦٓ",
"lemma": "تلاوت",
"word": 6
},
{
"ayah": "2:125",
"form": "مَثَابَةٗ",
"lemma": "مثابة",
"word": 4
},
{
"ayah": "2:125",
"form": "مُصَلّٗىۖ",
"lemma": "مصلى",
"word": 11
},
{
"ayah": "2:148",
"form": "وِجۡهَةٌ",
"lemma": "وجهة",
"word": 2
},
{
"ayah": "2:148",
"form": "مُوَلِّيهَاۖ",
"lemma": "مولي",
"word": 4
},
{
"ayah": "2:158",
"form": "ٱلصَّفَا",
"lemma": "صفا",
"word": 2
},
{
"ayah": "2:158",
"form": "وَٱلۡمَرۡوَةَ",
"lemma": "مروة",
"word": 3
},
{
"ayah": "2:158",
"form": "ٱعۡتَمَرَ",
"lemma": "اعتمر",
"word": 11
},
{
"ayah": "2:159",
"form": "ٱللَّـٰعِنُونَ",
"lemma": "لاعن",
"word": 20
},
{
"ayah": "2:164",
"form": "ٱلۡمُسَخَّرِ",
"lemma": "مسخر",
"word": 37
},
{
"ayah": "2:171",
"form": "يَنۡعِقُ",
"lemma": "ينعق",
"word": 6
},
{
"ayah": "2:177",
"form": "وَٱلۡمُوفُونَ",
"lemma": "موفي",
"word": 36
},
{
"ayah": "2:178",
"form": "ٱلۡقَتۡلَىۖ",
"lemma": "قتلى",
"word": 8
},
{
"ayah": "2:178",
Skeptic checks
- [neutral] Per morphology dataset quran-morphology-8f38b39 (QAC 0.4 derivative); rarity is annotation-dependent. Lemma normalization (letters-basic) merges spelling variants; counts shift under other normalizations.
- [weakens] Large hapax counts are a statistical property of essentially every natural-language corpus (Zipf tail); the list is useful for study, not remarkable in itself.
Your decision is recorded in data/research/reviews.json
420 roots appear in exactly one ayah unreviewed deterministic candidate
{
"count": 420,
"sample": [
{
"ayah": "2:16",
"root": "ربح",
"words": 1
},
{
"ayah": "2:61",
"root": "بصل",
"words": 1
},
{
"ayah": "2:61",
"root": "بقل",
"words": 1
},
{
"ayah": "2:61",
"root": "عدس",
"words": 1
},
{
"ayah": "2:61",
"root": "فوم",
"words": 1
},
{
"ayah": "2:61",
"root": "قثأ",
"words": 1
},
{
"ayah": "2:69",
"root": "فقع",
"words": 1
},
{
"ayah": "2:71",
"root": "وشي",
"words": 1
},
{
"ayah": "2:158",
"root": "مرو",
"words": 1
},
{
"ayah": "2:171",
"root": "نعق",
"words": 1
},
{
"ayah": "2:185",
"root": "رمض",
"words": 1
},
{
"ayah": "2:197",
"root": "زود",
"words": 2
},
{
"ayah": "2:255",
"root": "أود",
"words": 1
},
{
"ayah": "2:255",
"root": "وسن",
"words": 1
},
{
"ayah": "2:256",
"root": "فصم",
"words": 1
},
{
"ayah": "2:259",
"root": "سنه",
"words": 1
},
{
"ayah": "2:264",
"root": "صلد",
"words": 1
},
{
"ayah": "2:265",
"root": "طلل",
"words": 1
},
{
"ayah": "2:267",
"root": "غمض",
"words": 1
},
{
"ayah": "2:273",
"root": "لحف",
"words": 1
},
{
"ayah": "2:275",
"root": "خبط",
"words": 1
},
{
"ayah": "3:41",
"root": "رمز",
"words": 1
},
{
"ayah": "3:49",
"root": "ذخر",
"words": 1
},
{
"ayah": "3:61",
"root": "بهل",
"words": 1
},
{
"ayah": "3:75",
"root": "دنر",
"words": 1
},
{
"ayah": "3:156",
"root": "غزو",
"words": 1
},
{
"ayah": "3:159",
"root": "فظظ",
"words": 1
},
{
"ayah": "4:2",
"root": "حوب",
"words": 1
},
{
"ayah": "4:3",
"root": "عول",
"words": 1
},
{
"ayah": "4:6",
"root": "بدر",
"words": 1
},
{
"ayah": "4:21",
"root": "فضو",
"words": 1
},
{
"ayah": "4:51",
"root": "جبت",
"words": 1
},
{
"ayah": "4:56",
"root": "نضج",
"words": 1
},
{
"ayah": "4:71",
"root": "ثبي",
"words": 1
},
{
"ayah": "4:72",
"root": "بطأ",
"words": 1
},
{
"ayah": "4:83",
"root": "ذيع",
"words": 1
},
{
"ayah": "4:83",
"root": "نبط",
"words": 1
},
{
"ayah": "4:100",
"root": "رغم",
"words": 1
},
{
"ayah": "4:102",
"root": "سلح",
"words": 4
},
{
"ayah": "4:119",
"root": "بتك",
"words": 1
},
{
"ayah": "4:143",
"root": "ذبذب",
"words": 1
},
{
"ayah": "5:3",
"root": "خنق",
"words": 1
},
{
"ayah": "5:3",
"root": "ذكو",
"words": 1
},
{
"ayah": "5:3",
"root": "نطح",
"words": 1
},
{
"ayah": "5:3",
"root": "وقذ",
"words": 1
},
{
"ayah": "5:26",
"root": "تيه",
"words": 1
},
{
"ayah": "5:31",
"root": "بحث",
"words": 1
},
{
"ayah": "5:33",
"root": "نفي",
"words": 1
},
{
"ayah": "5:48",
"root": "نهج",
"words": 1
},
{
"ayah": "5:82",
"root": "قسس",
"words": 1
},
{
"ayah": "5:94",
"root": "رمح",
"words": 1
},
{
"ayah": "5:103",
"root": "سيب",
"words": 1
},
{
"ayah": "6:70",
"root": "بسل",
"words": 2
},
{
"ayah": "6:71",
"root": "حير",
"words": 1
},
{
"ayah": "6:95",
"root": "نوي",
"words": 1
},
{
"ayah": "6:99",
"root": "قنو",
"words": 1
},
{
"ayah": "6:99",
"root": "ينع",
"words": 1
},
{
"ayah": "6:143",
"root": "ضأن",
"words": 1
},
{
"ayah": "6:143",
"root": "معز",
"words": 1
},
{
"ayah": "6:146",
"root": "شحم",
"words": 1
},
{
"ayah": "7:18",
"root": "ذأم",
"words": 1
},
{
"ayah": "7:26",
"root": "ريش",
"words": 1
},
{
"ayah": "7:54",
"root": "حثث",
"words": 1
},
{
"ayah": "7:58",
"root": "نكد",
"words": 1
},
{
"ayah": "7:74",
"root": "سهل",
"words": 1
},
{
"ayah": "7:133",
"root": "ضفدع",
"words"Skeptic checks
- [neutral] Per morphology dataset quran-morphology-8f38b39 (QAC 0.4 derivative); rarity is annotation-dependent.
- [weakens] Single-context roots are expected in any corpus of this size; see the hapax note.
Your decision is recorded in data/research/reviews.json
Surahs with the highest and lowest root diversity per 100 words unreviewed deterministic candidate
{
"highest": [
{
"distinctRoots": 36,
"rootsPer100Words": 66.67,
"surah": 91,
"words": 54
},
{
"distinctRoots": 22,
"rootsPer100Words": 64.71,
"surah": 95,
"words": 34
},
{
"distinctRoots": 24,
"rootsPer100Words": 60,
"surah": 100,
"words": 40
},
{
"distinctRoots": 100,
"rootsPer100Words": 57.8,
"surah": 78,
"words": 173
},
{
"distinctRoots": 41,
"rootsPer100Words": 57.75,
"surah": 92,
"words": 71
}
],
"lowest": [
{
"distinctRoots": 477,
"rootsPer100Words": 14.37,
"surah": 7,
"words": 3320
},
{
"distinctRoots": 419,
"rootsPer100Words": 13.74,
"surah": 6,
"words": 3050
},
{
"distinctRoots": 442,
"rootsPer100Words": 12.7,
"surah": 3,
"words": 3481
},
{
"distinctRoots": 462,
"rootsPer100Words": 12.33,
"surah": 4,
"words": 3747
},
{
"distinctRoots": 587,
"rootsPer100Words": 9.6,
"surah": 2,
"words": 6116
}
],
"minimumWords": 30,
"surahsConsidered": 100
}Skeptic checks
- [weakens] Type diversity falls with text length by construction (roots repeat as texts grow); per-100-word normalization reduces but does not remove the confound. Comparisons across very different lengths remain unreliable.
- [neutral] Per morphology dataset quran-morphology-8f38b39 (QAC 0.4 derivative); rarity is annotation-dependent.
Your decision is recorded in data/research/reviews.json
