This is the label set and training recipe behind the proxy_v4 /
proxy_s2s3_v1 model card:
https://huggingface.co/LimeGS/herculaneum-legibility-proxy.
Read that first — it documents the model, the preprocessing, the
thresholds, and the validation results. This dataset only covers how the
model was trained: the human labeling protocol, the exact windows it
produced, and a script to turn them back into pixels… See the full description on the dataset page:
https://huggingface.co/datasets/LimeGS/herculaneum-legibility-proxy-labels.