RWKV World Corpus
(includes v3, v2.1 and v2 subsets)
This is an itemised and annotated list of the RWKV World corpus as described in the RWKV-7 paper
which is a multilingual dataset with about 3.1T tokens used to train the
"Goose" RWKV-7 World model series.
RWKV World v3 was crafted from public datasets spanning >100 world languages
(80% English, 10% multilang, and 10% code).