The Aquila-135M model is a small bilingual(Chinese and English) language model, which is trained using a two-phrase paradigm: pre-training and annealing.
This model used 1.66TB bilingual tokens in Chinese and English during pre-training phrase and 100B tokens during annealing training phrase.
In annealing stage, we selected 100B tokens of high-quality bilingual data and finally got our model.
The Aquila-135M-Instuct model is finetuned using Infinity Instruct.
The… See the full description on the dataset page:
https://huggingface.co/datasets/BAAI/Aquila-135M-Datasets.