Single-label classification: Use argmax on the logits
Multi-label / topic tagging: Take all categories above a confidence threshold (e.g. 0.15)
Confidence filtering: Discard predictions below 0.5 confidence for higher precision
Limitations
Trained on English-language news articles only — performance on non-news text or other languages will be lower
Labels were generated by GPT-5-nano, not human-annotated, so some label noise exists
Categories with fewer training examples (Holidays, Communication, Careers) may have lower accuracy
Very long articles are truncated to 1,024 tokens — classification is based on the beginning of the article
Citation
bibtex
1@misc{modernbert,
2 title={Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference},
3 author={Benjamin Warner and Antoine Chaffin and Benjamin Clavié and Orion Weller and Oskar Hallström and Said Taghadouini and Alexis Gallagher and Raja Biswas and Faisal Ladhak and Tom Aarsen and Nathan Cooper and Griffin Adams and Jeremy Howard and Iacopo Poli},
4 year={2024},
5 eprint={2412.13663},
6 archivePrefix={arXiv},
7}