Source: HRCS 2014, 2018, and 2022 direct award datasets.
Quality Filtering:
Only human-coded abstracts were included.
Records with abstracts shorter than 75 characters were removed during preprocessing to ensure the model had sufficient text to learn from.
Train/Test Split: The Test Set was isolated using only 2022 data to provide a modern performance benchmark.
To prevent the model from over-fitting on… See the full description on the dataset page:
https://huggingface.co/datasets/NIHRDataInsights/HRCSData.