v0.2.0 — supersedes the earlier data/wikipedia_ar/ v0.1.0 partial upload
(484 Arabic Wikipedia documents only). This release expands to the full 5-source
corpus below and moves the data to data/full_corpus/.
A multi-source Arabic/English text corpus about Palestinian history, culture, and
heritage, built for the Palestinian Cultural Knowledge
Platform
— a RAG + knowledge-graph research project. 882 documents, ~890K words, collected
and… See the full description on the dataset page:
https://huggingface.co/datasets/palestinian-kg/palestinian-cultural-knowledge.