Each line is one OCRFragment, a visible text region in one keyframe.
Common fields:
fragment_id: stable OCR observation identifier
video_id: source video identifier
time_span: [timestamp_ms, timestamp_ms] in source video time
text: recognized visible text
bbox: normalized text region with x_min, y_min, x_max, and y_max
There is no placeholder for frames without text and no fabricated confidence score. Join OCR with… See the full description on the dataset page:
https://huggingface.co/datasets/egolqa/ocrv2.