This is a multiturn instruct tuning dataset with 2,333,924 trainable tokens, created with Augmentoolkit, covering the material in the majority of the US Army Field Manuals that are publicly available.
Unlike many previous Augmentoolkit datasets, the questions and answers here are without fluff and are more "to the point". This "sharper" data is intended to help the LLM with recalling facts.
There are three main datasets included here: "vanilla", "negative" and "long".
Vanilla data is simple… See the full description on the dataset page:
https://huggingface.co/datasets/BadBerad5222/us-army-fm-instruct.