This dataset contains definitions and important facts about 467k words that appear in the context of English texts.
It has been used to train our high-performance, compact text embedding models mdbr-leaf-ir and mdbr-leaf-mt.
The original list of words stems from here. We have extended it with definitions and important facts about each word using Claude 3.7 Sonnet.