The Winograd schema challenge composes tasks with syntactic ambiguity,
which can be resolved with logic and reasoning (Levesque et al., 2012).
The texts for the Winograd schema problem are obtained using a semi-automatic
pipeline. First, lists of 11 typical grammatical structures with syntactic
homonymy (mainly case) are compiled. For example, two noun phrases with a
complex subordinate: 'A trinket from Pompeii that has survived the centuries'.
Requests corresponding to these constructions are submitted in search of the
Russian National Corpus, or rather its sub-corpus with removed homonymy. In the
resulting 2+k examples, homonymy is removed automatically with manual validation
afterward. Each original sentence is split into multiple examples in the binary
classification format, indicating whether the homonymy is resolved correctly or
not.