Stratified random split with fully separated sets:
Train (80%): 30,806,432 rows — model training and architecture search
Test (10%): 3,850,804 rows — threshold calibration (held out from training)
Validation (10%): 3,850,805 rows — final reported metrics (never touched)