This dataset contains the annotated TLUnified corpora from Cruz and Cheng
(2021). It is a curated sample of around 7,000 documents for the named entity
recognition (NER) task. The majority of the corpus are news reports in Tagalog,
resembling the domain of the original ConLL 2003. There are three entity types:
Person (PER), Organization (ORG), and Location (LOC).