This dataset is a preprocessed version of microsoft/FStarDataSet-V2. It has been reformatted into a chat-style JSONL structure for supervised fine-tuning of language models on F* function synthesis and proof completion.
Each line in these files is a JSON object with the following schema (where the keys correspond to… See the full description on the dataset page:
https://huggingface.co/datasets/dassarthak18/FStarDataset-V2-Conversation.