Phoneme-based GPT-2 models trained on all 31 sections of the
IPA-CHILDES dataset for the paper
BabyLM's First Words: Word Segmentation as a Phonological Probing Task.
The models have 600k non-embedding parameters and were trained on 100k tokens of their language. They were evaluated for phonological knowledge using the
word segmentation task. Check out the paper for more details. Training and analysis scripts can be found
here.
1from transformers import AutoModel
2farsi_model = AutoModel.from_pretrained('phonemetransformers/ipa-childes-models-tiny', subfolder='Farsi')