This is a training dataset for paper "DuDi: Dual-Signal Distillation with Cross-Lingual Verbalizer
". We use SEA-Instruct, which covers seven SEA languages: Indonesian, Vietnamese, Thai, Tamil, Tagalog, Malay, and Burmese.
The dataset contains open-source prompts, each paired with a synthetic response and quality estimate.
We sample 4,000 high-quality examples per language, as labeled by the original dataset, resulting 28,000 samples.
Random sampling constraints preserve the distribution of… See the full description on the dataset page:
https://huggingface.co/datasets/aisingapore/Small-SEA-Instruct-2602.