This repository contains the unified, high-density raw text data used for the Continued Pre-Training (CPT) of the Sensix Paite models (Gemma-4-31B Master and Gemma-4-2B/5B Nitro).
The dataset is specifically designed to expand a base model's vocabulary and internalize Paite linguistic patterns, syntax, and tonal logic before moving to instruction fine-tuning (SFT).
This is a unified dataset… See the full description on the dataset page:
https://huggingface.co/datasets/sensix-zo/Continued-Pre-Training-CPT-Paite.