Dataset contains articles from Wikipedia Bahasa Indonesia which fulfill these conditions:
The pages contain many noun phrases, which the authors subjectively pick: (i) fictional plots, e.g., subtitles for films,
TV show episodes, and novel stories; (ii) biographies (incl. fictional characters); and (iii) historical events or important events.
The pages contain significant variation of pronoun and named-entity. We count the number of first, second, third person pronouns,
and clitic pronouns in the document by applying string matching.We examine the number
of named-entity using the Stanford CoreNLP
NER Tagger (Manning et al., 2014) with a
model trained from the Indonesian corpus
taken from Alfina et al. (2016).
The Wikipedia texts have length of 500 to
2000 words.
We sample 201 of pages from subset of filtered
Wikipedia pages. We hire five annotators who are
undergraduate student in Linguistics department.
They are native in Indonesian. Annotation is carried out using the Script d’Annotation des Chanes
de Rfrence (SACR), a web-based Coreference resolution annotation tool developed by Oberle (2018).
From the 201 texts, there are 16,460 mentions
tagged by the annotators