This dataset contains vi subsets (first 191 examples) and auto-translation from en to vi subsets (the rest, 38346 examples) from OASST1. All auto-translation examples are generated using VietAI envit5-translation.
The vi subsets have the same features as the original dataset. Meanwhile, the auto-translation subsets introduce two new features:
"text_chunks" is a list that contains chunked text split from "text", each chunk has no more than 300 tokens. The sent_tokenizer and word_tokenzier… See the full description on the dataset page:
https://huggingface.co/datasets/Zayt/oasst1-vi.