CrawList is a dataset of URLs to crawl for extra information in multiple tasks, such as retrieval of public documentation or grounding LLMs on the latest data existing on the internet.
Dataset Details
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More Information Needed]
Out-of-Scope Use… See the full description on the dataset page: https://huggingface.co/datasets/dalitics/CrawList.