Gold-standard labelled job postings sampled daily from public company
career pages. Produced by a Codex-first agent pipeline with task-specific
subagents for HTML normalization, section splitting, and structured
extraction. The dataset is the substrate for training an improved
structured-information extractor for jseek.co.
Current row counts by date: 2026-08-25: 10 · 2026-08-24: 10 · 2026-08-23: 10 · 2026-08-22: 10 · 2026-08-21: 10 · 2026-08-20: 10 ·… See the full description on the dataset page:
https://huggingface.co/datasets/viktor-shcherb/jobseek-postings-labelled.