A large-scale, labeled dataset of 76,000+ images extracted from U.S. IPO registration statements (S-1 and F-1 filings) filed with the SEC EDGAR system, spanning 1994–2026.
Every image has been classified through a multi-stage pipeline: initial detection with YOLOv8, followed by verification from an ensemble of 8 Vision-Language Models (VLMs). Chart images include additional structured metadata describing chart type, visual properties, and content.… See the full description on the dataset page:
https://huggingface.co/datasets/gtfintechlab/ipo-images.