ARXIV:2603.23972 · LLM GROUNDING · SUBMITTED 26 MAR · 20:30 UTC · FRESHNESS STALE

VerifiedSource: PDF linkedVerifiedPaperPack: citation fields availablePartialProof: partial proof status

Grounding Arabic LLMs in the Doha Historical Dictionary: Retrieval-Augmented Understanding of Quran and Hadith

Somaya Eltanbouly · Samer Rashwani · arXiv

A retrieval-augmented generation framework that grounds Arabic LLMs in historical lexicographic data to significantly improve understanding of religious texts.

Ship in 2-4 weeks›Score7.0Evidence partial

Opportunity summary

Pain A retrieval-augmented generation framework that grounds Arabic LLMs in historical lexicographic data to significantly improve understanding of religious texts.

Evidence 0 refs | 0 sources | 50% coverage

Blocker Evidence partial

Open Build Read PDF Signal Canvas Track

PROBLEM

A retrieval-augmented generation framework that grounds Arabic LLMs in historical lexicographic data to significantly improve understanding of religious texts. To address this limitation, we develop a retrieval-augmented generation (RAG) framework grounded in diachronic lexicographic…

METHOD

Full abstract

Large language models (LLMs) have achieved remarkable progress in many language tasks, yet they continue to struggle with complex historical and religious Arabic texts such as the Quran and Hadith. To address this limitation, we develop a retrieval-augmented generation (RAG) framework grounded in diachronic lexicographic knowledge. Unlike prior RAG systems that rely on general-purpose corpora, our approach retrieves evidence from the Doha Historical Dictionary of Arabic (DHDA), a large-scale resource documenting the historical development of Arabic vocabulary. The proposed pipeline combines hybrid retrieval with an intent-based routing mechanism to provide LLMs with precise, contextually relevant historical information. Our experiments show that this approach improves the accuracy of Arabic-native LLMs, including Fanar and ALLaM, to over 85\%, substantially reducing the performance gap with Gemini, a proprietary large-scale model. Gemini also serves as an LLM-as-a-judge system for automatic evaluation in our experiments. The automated judgments were verified through human evaluation, demonstrating high agreement (kappa = 0.87). An error analysis further highlights key linguistic challenges, including diacritics and compound expressions. These findings demonstrate the value of integrating diachronic lexicographic resources into retrieval-augmented generation frameworks to enhance Arabic language understanding, particularly for historical and religious texts. The code and resources are publicly available at: https://github.com/somayaeltanbouly/Doha-Dictionary-RAG.

RESULT

ScienceToStartup currently rates this 7.0/10 on the public viability pass. Our experiments show that this approach improves the accuracy of Arabic-native LLMs, including Fanar and ALLaM, to over 85\%, substantially reducing the performance gap…

WHY NOW

LLM Grounding moved forward this cycle; last verified April 2026. Public score 7.0/10. Implementation evidence is present through a linked repository.

Continue into Read for claims, analysis, references, and neighboring papers.

Opportunity summary

Score7.0

PainA retrieval-augmented generation framework that grounds Arabic LLMs in historical lexicographic data to significantly improve understanding of religious texts.

Evidence0 refs | 0 sources | 50% coverage

Blockerno shell-level blocker reported

Analysis summary

A retrieval-augmented generation framework that grounds Arabic LLMs in historical lexicographic data to significantly improve understanding of religious texts.

VerifiedSource: PDF linkedVerifiedPaperPack: citation fields availablePartialProof: partial proof status

Competitive landscape

A retrieval-augmented generation framework that grounds Arabic LLMs in historical lexicographic data to significantly improve understanding of religious texts.

Segment

LLM Grounding

Adoption evidence

Public code linked for build inspection

Commercial read

7.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

{ "contract_version": "paper-r2", "paper_id": "fbefeecf-df4c-44a8-b6cb-e32ddd1317c7", "arxiv_id": "2603.23972", "canonical_route": "/paper/grounding-arabic-llms-in-the-doha-historical-dictionary-retrieval-augmented-understanding-of-quran-and-hadith", "active_tab": "synced from current hash by the drawer client", "selected_artifact": "grounding-arabic-llms-in-the-doha-historical-dictionary-retrieval-augmented-understanding-of-quran-and-hadith", "endpoints": { "paper_pack": "/api/v1/paper/grounding-arabic-llms-in-the-doha-historical-dictionary-retrieval-augmented-understanding-of-quran-and-hadith/paper-pack", "build_passport": "/api/v1/paper/grounding-arabic-llms-in-the-doha-historical-dictionary-retrieval-augmented-understanding-of-quran-and-hadith/build-passport", "mcp_resource": "sciencetostartup://surfaces/paper-workspace" } }

{ "surface": "paper", "mode": "paper", "query": "Grounding Arabic LLMs in the Doha Historical Dictionary: Retrieval-Augmented Understanding of Quran and Hadith", "normalized_query": "2603.23972", "route": "/paper/grounding-arabic-llms-in-the-doha-historical-dictionary-retrieval-augmented-understanding-of-quran-and-hadith", "paper_ref": "grounding-arabic-llms-in-the-doha-historical-dictionary-retrieval-augmented-understanding-of-quran-and-hadith", "topic_slug": null, "benchmark_ref": null, "dataset_ref": null }

{ "@context": "https://schema.org", "@graph": [ { "@type": "WebPage", "@id": "https://sciencetostartup.com/paper/grounding-arabic-llms-in-the-doha-historical-dictionary-retrieval-augmented-understanding-of-quran-and-hadith#webpage", "url": "https://sciencetostartup.com/paper/grounding-arabic-llms-in-the-doha-historical-dictionary-retrieval-augmented-understanding-of-quran-and-hadith", "name": "Grounding Arabic LLMs in the Doha Historical Dictionary: Retrieval-Augmented Understanding of Quran and Hadith", "description": "A retrieval-augmented generation framework that grounds Arabic LLMs in historical lexicographic data to significantly improve understanding of religious texts.", "isPartOf": { "@id": "https://sciencetostartup.com/#website" } }, { "@type": "ScholarlyArticle", "@id": "https://sciencetostartup.com/paper/grounding-arabic-llms-in-the-doha-historical-dictionary-retrieval-augmented-understanding-of-quran-and-hadith#scholarlyArticle", "headline": "Grounding Arabic LLMs in the Doha Historical Dictionary: Retrieval-Augmented Understanding of Quran and Hadith", "description": "A retrieval-augmented generation framework that grounds Arabic LLMs in historical lexicographic data to significantly improve understanding of religious texts.", "url": "https://sciencetostartup.com/paper/grounding-arabic-llms-in-the-doha-historical-dictionary-retrieval-augmented-understanding-of-quran-and-hadith", "sameAs": "https://arxiv.org/abs/2603.23972", "identifier": { "@type": "PropertyValue", "propertyID": "arXiv", "value": "2603.23972" }, "isAccessibleForFree": true, "isPartOf": { "@id": "https://sciencetostartup.com/#website" }, "datePublished": "2026-03-25T06:09:42.000Z", "author": [ { "@type": "Person", "name": "Somaya Eltanbouly" }, { "@type": "Person", "name": "Samer Rashwani" } ], "codeRepository": "https://github.com/somayaeltanbouly/Doha-Dictionary-RAG", "additionalProperty": [ { "@type": "PropertyValue", "propertyID": "viabilityScore", "value": 7 }, { "@type": "PropertyValue", "propertyID": "researchDomain", "value": "LLM Grounding" }, { "@type": "PropertyValue", "propertyID": "commercialReadiness", "value": "code, repo url" } ] }, { "@type": "SoftwareSourceCode", "@id": "https://sciencetostartup.com/paper/grounding-arabic-llms-in-the-doha-historical-dictionary-retrieval-augmented-understanding-of-quran-and-hadith#software", "name": "Grounding Arabic LLMs in the Doha Historical Dictionary: Retrieval-Augmented Understanding of Quran and Hadith - Source Code", "description": "A retrieval-augmented generation framework that grounds Arabic LLMs in historical lexicographic data to significantly improve understanding of religious texts.", "codeRepository": "https://github.com/somayaeltanbouly/Doha-Dictionary-RAG", "url": "https://github.com/somayaeltanbouly/Doha-Dictionary-RAG" }, { "@type": "BreadcrumbList", "itemListElement": [ { "@type": "ListItem", "position": 1, "name": "Home", "item": "https://sciencetostartup.com" }, { "@type": "ListItem", "position": 2, "name": "LLM Grounding", "item": "https://sciencetostartup.com/topics" }, { "@type": "ListItem", "position": 3, "name": "Grounding Arabic LLMs in the Doha Historical Dictionary: Ret", "item": "https://sciencetostartup.com/paper/grounding-arabic-llms-in-the-doha-historical-dictionary-retrieval-augmented-understanding-of-quran-and-hadith" } ] } ] }

Competitive landscape

A retrieval-augmented generation framework that grounds Arabic LLMs in historical lexicographic data to significantly improve understanding of religious texts.

Segment

LLM Grounding

Adoption evidence

Public code linked for build inspection

Commercial read

7.0/10 public viability

Direct

not classified

Adjacent

not classified

Substitute

not classified

Unknown

not classified

Grounding Arabic LLMs in the Doha Historical Dictionary: Retrieval-Augmented Understanding of Quran and Hadith

Grounding Arabic LLMs in the Doha Historical Dictionary: Retrieval-Augmented Understanding of Quran and Hadith

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline

Claim map

Constellation map

Competitive landscape

Buzz

PDF

REFERENCES

Related Papers

Subscribe to the weekly brief

Build artifacts

Brief

Experiment plan

Validation checklist

Scientific founder

Translational engineer

Domain operator

GTM lead

Regulatory/clinical advisor

Timeline