Source: cminst/transcoda-synth-300k-prepared-trainval-v2 train + cminst/transcoda-rare-situation-oversample-v1/rare_situation_weighted_train_manifest.jsonl
This is a canonical standardized Transcoda dataset. All published splits use the standard Hugging Face Datasets Parquet layout under data/.
Canonical columns:
image
transcription
sample_id
source
metadata: JSON string preserving source/provenance fields
train: 243113 rows, 244 Parquet… See the full description on the dataset page:
https://huggingface.co/datasets/cminst/transcoda-rare-oversample-243k.