A dataset of texts paired with coarse rewrites of themselves — a
lower-resolution version of the text (roughly 1/4 the original length), not a
summary about it. Generated from Skylion007/openwebtext
using Qwen/Qwen3-4B-Instruct-2507.
This release covers the first 1,000,000 documents of OpenWebText (sequential,
rows 0–999,999). The repo is intended to grow toward the full corpus in future
releases.
column
type
description… See the full description on the dataset page:
https://huggingface.co/datasets/EER6/openwebtext-coarse.