We release 44.4B tokens of high-quality, model-filtered synthetic texts obtained via our REcycling the Web with guIded REwrite (REWIRE) approach.
The generation process involves taking all documents that are of moderate quality (i.e., having passed some rule-based filters),
using an LLM (Llama-3.3-70B-Instruct) to identify the purpose of the text content, and then asking the LLM to come up with an improved document conditioned on… See the full description on the dataset page:
https://huggingface.co/datasets/facebook/recycling_the_web.