PL-ModernBERT-wp-es is a phoneme-level masked language model trained on Spanish text. It is based on the PL-BERT architecture, which learns phoneme representations via a BERT-style masked language modeling objective.
This model is designed to support phoneme-based text-to-speech (TTS) systems, including but not limited to StyleTTS2.
Features of our PL-ModernBERT:
It is trained exclusively on Spanish phonemized text
It uses a reduced phoneme vocabulary of 178 tokens
It uses wordpiece tokenizer
It includes a custom token_maps.pkl and adapted util.py
Intended Uses and Limitations
Intended uses
Integration into phoneme-based TTS pipelines such as StyleTTS2, Matxa-TTS, or custom diffusion-based synthesizers
Accent-aware synthesis and phoneme embedding extraction for Catalan
Research on phoneme-level language modeling in low-resource or multi-accent settings
Limitations
Not designed for general NLP tasks like classification or sentiment analysis
Only supports Spanish phoneme tokens
Some accents may be underrepresented in the training data
How to Get Started with the Model
Here is an example of how to use this model within the StyleTTS2 framework:
Inside the Utils directory, create a new folder, for example: PLBERT_es_
Copy the following files into that folder:
config.yml (training configuration)
step_1000000.t7 (trained checkpoint)
token_maps.pkl (phoneme to ID mapping)
util.py (modified to fix position ID loading)
In your StyleTTS2 configuration file, update the PLBERT_dir entry to:
PLBERT_dir: Utils/PLBERT_es
Update the import statement in your code to:
from Utils.PLBERT_es.util import load_plbert
Note: Although this example uses StyleTTS2, the model is compatible with other TTS architectures that operate on phoneme sequences. You can use the contextualized phoneme embeddings from PL-BERT in any compatible synthesis system.
Training Details
Training data
The model was trained on a Spanish corpus phonemized (any phonemizer can be used). It uses a consistent phoneme token set with boundary markers and masking tokens.
Tokenizer: wordpiece
Phoneme masking strategy: word-level and phoneme-level masking and replacement
Training steps: 1,000,000
Precision: Mixed (fp16)
Training configuration
Model parameters:
Vocabulary size: 178
Hidden size: 768
Attention heads: 12
Intermediate size: 2048
Number of layers: 12
Max position embeddings: 512
Dropout: 0.1
Other parameters:
Batch size: 8
Max mel length: 512
Word mask probability: 0.15
Phoneme mask probability: 0.1
Replacement probability: 0.2
Token separator: space
Token mask: M
Word separator ID: 102
Evaluation
The model has not been benchmarked via perplexity or extrinsic evaluation, but has been successfully integrated into TTS pipelines such as StyleTTS2, where it enables the synthesis of Spanish.
Citation
If this code contributes to your research, please cite the work:
@misc{zevallos2025plmodernbertwpes,
title={PL-ModernBERT-wp-es},
author={Rodolfo Zevallos, Jose Giraldo and Carme Armentano-Oller},
organization={Barcelona Supercomputing Center},
url={https://huggingface.co/BSC-LT/PL-ModernBERT-wp-es},
year={2025}
}
This work is funded by the Ministerio para la Transformación Digital y de la Función Pública - Funded by EU – NextGenerationEU within the framework of the project Desarrollo de Modelos ALIA.