A 90K-example subset of the full JaViCorpus ChatML translation dataset, optimized for fast SFT training iterations on Google Colab (T4 GPU).
For the full ~480K dataset, see tranguyenxuwu/javicorpus-chatml-translation.
This is a strategically sampled subset of the bidirectional Japanese↔Vietnamese translation dataset derived from the ngovinhtn/JaViCorpus parallel corpus collection. Each example is a 3-turn… See the full description on the dataset page:
https://huggingface.co/datasets/tranguyenxuwu/javicorpus-chatml-translation-mini.