Viewer-friendly Persian / Arabic-script subset extracted from the current ultra-clean UncGPT training candidate corpus for manual inspection before training.
This dataset is for review/QC. It includes conversations selected when either:
the language hygiene audit marked the conversation as fa, or
Arabic-script characters appeared anywhere in seeker/Uncle text.
conversations: one row per conversation, with… See the full description on the dataset page:
https://huggingface.co/datasets/Reza2kn/uncgpt-persian-ultraclean-review-2026-05-14.