The following dataset is a collection of JSON sets that contain most of the Hebrew Wikipedia.
Crawl Hebrew Wikipedia: Begin by crawling through Hebrew Wikipedia to collect all redirect links on each page.
Breadth-First Search (BFS): For each page, apply a BFS-like strategy to ensure that every link is scraped.
Link Collection: Collect the links as… See the full description on the dataset page:
https://huggingface.co/datasets/YanFren/Hebrew_wikipedia.