Canonical minimal dataset for Japanese IME candidate reranking. Every row has
only context, reading, and target.
core: 8,084
anonymized person context: 219
synthetic hard context: 146,270
difficult context: 11,329
person_context replaces concrete people and public persona names with
[PERSON]. hard_context is kept separate so callers can choose its sampling
weight rather than allowing a large synthetic set to dominate ordinary… See the full description on the dataset page:
https://huggingface.co/datasets/fa0311/warabi-context-dataset.