Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC),
used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v1.
IRC full text stored in a Qdrant vector store (chunked at ~512 tokens)
An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk
Generated pairs are deduplicated and split into train/validation
Split
Records… See the full description on the dataset page:
https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-36k.