Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
47,679 chunks from 13,572 source documents
38,603 Vietnamese chunks and 9,076 English chunks
Maximum chunk length: 512 BGE-M3 tokenizer tokens
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page:
https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked.