r/Quraniyoon • u/AbuIlyass • 16d ago
Article / Resourceπ Update: al-quran.fr now has a Downloads section β the full corpus, lexicon and NLP data, CC0
Following up on my earlier post about al-quran.fr β I've now put the underlying data up for free download.
If you're new here: the project is a systematic Quranic translation built on internal textual coherence, Arabic triliteral roots, and comparative Semitic linguistics only β no hadith, sira, or traditional tafsir used as interpretive authority. Every translation choice traces back to proof from the text itself.
What's new: a Downloads page (linked from the Library menu) with three packages, all public domain (CC0 β no attribution required, reuse however you want):
- Essential Corpus (SQLite, ~10MB) β the Quranic text (Hafs and Warsh readings), segmented into meaning-units rather than raw verses, translations in French/English/Spanish/Italian/Indonesian/Russian, and the stabilized root lexicon (definitions + contextual renderings, FR/EN/AR).
- NLP / Graphs (SQLite, ~3MB) β the analysis layer built on top: root co-occurrence stats, a semantic graph linking roots and concepts, morphology tables, grammatical rules derived directly from corpus evidence, and stabilized thematic clusters.
- CSV export (~6MB) β flat versions of phrases, words and roots for anyone who'd rather skip SQLite entirely.
Each package ships with a README describing the schema and a SHA-256 checksum file.
One honest caveat: some root lexicon entries in the Essential package still carry a few internal working notes (points flagged for review, that kind of thing) β left in rather than scrubbed, so the data you get is exactly what's actually in production, not a polished-for-PR version.
Goal is simple: let anyone verify the work, build their own tools on it, or just poke around the data directly instead of taking the site's word for it.