This is a lengthened version of Mutonix/RefGPT-Fact-v2. The reference field in each sample is 8 times the length of the original sample. Correspondingly, the dataset is subsampled to 1/8 of its original size.
Each reference is upsampled by K times, where:
K∼Poisson(λ=8)
K \sim \text{Poisson}(\lambda=8)
K∼Poisson(λ=8)
To lengthen the reference text, the reference text is shuffled into K - 1 paragraphs. The K - 1… See the full description on the dataset page:
https://huggingface.co/datasets/wenbopan/RefGPT-Fact-v2-8x.