For non-table text inputs, the target is a JSON list of numeric entity/datatype pairs enriched with the sentence containing the matched numeric entity.
The dataset is derived from FinTagging_800_200_HF and preserves the original
train/test assignment by source_sample_idx and context_id. The XBRL concept
tag is intentionally omitted from the target.
Split
Samples
Output entries
Duplicate output key… See the full description on the dataset page:
https://huggingface.co/datasets/lm2445/FinTagging1000_text.