This dataset is a large English-only Bluesky text corpus curated for tokenizer and language-model experiments.
The corpus was built in two stages:
Stage 1: a full download of the Hugging Face dataset Roronotalt/bluesky, filtered to English posts and deduplicated by post URI.
Stage 2: additional English Bluesky posts collected directly from Bluesky repositories using a custom extractor, again deduplicated by URI and checked… See the full description on the dataset page: https://huggingface.co/datasets/akarrouch-mohamed/bluesky.