💻Github Repo 🖨️arXiv Paper
The official pre-training dataset for "Towards Building Multilingual Language Model for Medicine".
We add Arabic and German corpus to MMedC.
This repo contains MMedC, a multilingual medical corpus with 25.5 billion tokens.
Spanish
Indo-European
3.98
0.31
0.05
0.02… See the full description on the dataset page:
https://huggingface.co/datasets/22yuan/MMedC.