Views
No views yet
| Domain | Dataset | Data Size | #Docs | #Tokens |
|---|---|---|---|---|
| Formal | Wikipedia | 9GB | 2,665,357 | 1.9B |
| Formal | News | 28GB | 12,305,326 | 6.1B |
| Formal | GC4 | 90GB | 31,669,772 | 19.4B |
| Informal | Reddit 2019-2023 (GER) | 5.8GB | 15,036,592 | 1.3B |
| Informal | Holiday Reviews | 2GB | 4,876,405 | 428M |
| Legal | OpenLegalData: German cases and laws | 5.4GB | 308,228 | 1B |
| Medical | Smaller public datasets | 253MB | 179,776 | 50M |
| Medical | CC medical texts | 3.6GB | 2,000,000 | 682M |
| Medical | Medicine Dissertations | 1.4GB | 14,496 | 295M |
| Medical | Pubmed abstracts (translated) | 8.5GB | 21,044,382 | 1.7B |
| Medical | MIMIC III (translated) | 2.6GB | 24,221,834 | 695M |
| Medical | PMC-Patients-ReCDS (translated) | 2.1GB | 1,743,344 | 414M |
| Literature | German Fiction | 1.1GB | 3,219 | 243M |
| Literature | English books (translated) | 7.1GB | 11,038 | 1.6B |
| - | Total | 167GB | 116,079,769 | 35.8B |
| Model | GE14 | GQuAD | GE18 | TS | GGP | GRAS1 | JS | DROC | Avg |
|---|---|---|---|---|---|---|---|---|---|
| GBERTbase | 87.10±0.12 | 72.19±0.82 | 51.27±1.4 | 72.34±0.48 | 78.17±0.25 | 62.90±0.01 | 77.18±3.34 | 88.03±0.20 | 73.65±0.50 |
| GELECTRAbase | 86.19±0.5 | 74.09±0.70 | 48.02±1.80 | 70.62±0.44 | 77.53±0.11 | 65.97±0.01 | 71.17±2.94 | 88.06±0.37 | 72.71±0.66 |
| GottBERT | 87.15±0.19 | 72.76±0.378 | 51.12±1.20 | 74.25±0.80 | 78.18±0.11 | 65.71±0.01 | 74.60±4.75 | 88.61±0.23 | 74.05±0.51 |
| GeBERTabase | 88.06±0.22 | 78.54±0.32 | 53.16±1.39 | 74.83±0.36 | 78.13±0.15 | 68.37±1.11 | 81.85±5.23 | 89.14±0.32 | 76.51±0.32 |
1@inproceedings{dada2023impact,
2 title={On the Impact of Cross-Domain Data on German Language Models},
3 author={Dada, Amin and Chen, Aokun and Peng, Cheng and Smith, Kaleb E and Idrissi-Yaghir, Ahmad and Seibold, Constantin Marc and Li, Jianning and Heiliger, Lars and Friedrich, Christoph M and Truhn, Daniel and others},
4 booktitle={The 2023 Conference on Empirical Methods in Natural Language Processing},
5 year={2023}
6}