Benchmark for the US tax code. Train set contains 10,833 questions and test set contains 1,204 questions. The questions are all multiple choice.
All questions were generated by o3-high using the full US tax code as a source.
Model
Without RAG
With RAG
gpt-4.1-nano
72.67 %
75.42 %
gpt-4.1-mini
78.90 %
86.30 %
o4-mini-medium
82.97 %
89.37 %
gpt-4.1
86.05 %
85.88 %
Claude 4 Sonnet (non-thinking)
90.53 %
Untested
gpt-4.5-preview
91.36 %… See the full description on the dataset page:
https://huggingface.co/datasets/trentmkelly/USTaxCodeBench.