Two document-to-LoRA (doc2lora) continual-learning pretraining datasets, sharing the same
p0_pwc_compact shard layout (train/*.parquet + manifest.json).
Subfolders
doc2lora_full_p0p1_text_v1/
Plain-text fields (context/prompt/response/suffix) plus Qwen3.5-4B token id lists
(*_tokens_qwen35) for arbitrary-model re-tokenization.
3287 parquet shards, ~7.7 GB.
doc2lora_full_p0p1_ce_qwen35tok_v1/… See the full description on the dataset page: https://huggingface.co/datasets/Hizy/doc2qwen.