This is a randomly split five fold dataset of the Pretrain MIMIC dataset stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is pretrain-mimic-250-1250 where 250 refers to the sampling frequency and 1250 denotes 5 seconds.
The code to do the splitting is here.
Any questions or issues, please do not hesitate to reach out to the maintainer of ECG-Bench.