This dataset is a subset of the original nearData dataset, prepared for the continued pre-training of a pre-trained LLM.
The idea behind the continued pre-training of pre-trained models is to further instruct them with specific information, in this case on the Near Protocol blockchain, before fine-tuning them.
The preTrainingNEAR dataset was prepared from local text files using the datasets library from Hugging Face. It includes:
nearBlog: 481 blog articles from Near Blog collected on March… See the full description on the dataset page:
https://huggingface.co/datasets/jcarbonnell/preTrainingNEAR.