SynthFormulaNet is a multimodal dataset designed for training the SmolDocling model. It contains over 6.4 million pairs of synthetically rendered images depicting mathematical formulas and their corresponding LaTeX representations. The LaTeX data was collected from permissively licensed sources, and the images were generated using LaTeX at 120 DPI with diverse rendering styles, fonts, and layout configurations to maximize visual variability. This dataset also… See the full description on the dataset page:
https://huggingface.co/datasets/docling-project/SynthFormulaNet.