A 20,300-row citation-grounded synthetic fraud-narrative dataset for four
underserved US financial-system archetypes — remittance, gig_worker,
unbanked, ITIN — with all 25 FinCEN typology codes exercised.
V3 covered 10 of 25 FinCEN typology codes. v4 closed the gap to 18/25
through three targeted changes:
16 persona edits documenting fraud events (SIM-swap, BEC, hawala/IVTS… See the full description on the dataset page:
https://huggingface.co/datasets/Nachammai41/underserved-persona_conditioned-fraud-v4.