Tokenized English-Turkish Translation Corpus
Format-augmented and tokenized subset of Ba2han/finetranslations-TR_filtered, built with
Ba2han/outputs.
Metric
Amount
Target tokens
500,000,000
Actual tokens
500,000,186
Full-example overshoot
186
Output rows
343,740
Average tokens/row
1,454.59
Minimum tokens/row
441
Maximum tokens/row
5,472
Source rows scanned
204,362
Source rows passing filters
42,976
Parquet shards
4… See the full description on the dataset page:
https://huggingface.co/datasets/Ba2han/tokenized_translation.