A 1-billion-token representative sample of a much larger cleaned Australian web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. π
The full 294B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. π Researchers seeking access to the full corpus forβ¦ See the full description on the dataset page:
https://huggingface.co/datasets/JoeyLLM/australian-dataset-1b.