Part-of-speech information is basic NLP task. However, Twitter text
is difficult to part-of-speech tag: it is noisy, with linguistic errors and idiosyncratic style.
This dataset contains two datasets for English PoS tagging for tweets:
- Ritter, with train/dev/test
- Foster, with dev/test
Splits defined in the Derczynski paper, but the data is from Ritter and Foster.
For more details see: