We introduce TexTAR, a multi-task, context-aware Transformer for Textual Attribute Recognition (TAR),
capable of handling both positional cues (bold, italic) and visual cues (underline, strikeout) in
noisy, multilingual document images.
MMTAD (Multilingual Multi-domain Textual Attribute Dataset) comprises 1,623 real-world document images—from legislative records and notices to textbooks and notary documents—captured under diverse lighting, layout, and noise… See the full description on the dataset page:
https://huggingface.co/datasets/Tex-TAR/MMTAD.