Source: cminst/transcoda-synth-300k-prepared-trainval-v2::transcoda_synth_300k_prepared_trainval_v2.tar.zst!transcoda_synth_public_300k
This is the canonical standardized Transcoda synthetic dataset. All published
splits use the standard Hugging Face Datasets Parquet layout under data/.
Canonical columns:
image
transcription
sample_id
source
metadata: JSON string preserving source/provenance fields
train: 243113 rows, 300 Parquet shards, 27.27 GiB… See the full description on the dataset page:
https://huggingface.co/datasets/cminst/transcoda-synth-243k.