UniRec-0.1B is a unified recognition model with only 0.1B parameters, designed for high-accuracy and efficient recognition of plain text (words, lines, paragraphs), mathematical formulas (single-line, multi-line), and mixed content in both Chinese and English.
It addresses structural variability and semantic… See the full description on the dataset page:
https://huggingface.co/datasets/topdu/UniRec40M.