A multi-domain English NLP corpus with four annotation layers: Named Entity Recognition (18 types), POS tagging (17 UPOS tags), dependency parsing, and dialog act classification (9 types). All data is commercially licensed (CC BY-SA 4.0 compatible).
Built for training kniv multi-task NLP models that power the uniko cognitive memory system.