This is the dataset for finetuning TF-ID models.
It contains about 4,600 images (academic paper pages) with bounding boxes of tables and figures in coco format.
The papers are selected from Hugging Face Daily Papers, covering mostly AI/ML/DL related topics.
You can use this dataset to reproduce all TF-ID models.
All bounding boxes were annotated manually by Yifei Hu
Unzip the… See the full description on the dataset page:
https://huggingface.co/datasets/yifeihu/TF-ID-arxiv-papers.