NewTerm is a benchmark dataset designed to evaluate the ability of large language models (LLMs) to understand new terms in real time. It addresses the less-explored area of new term evaluation by providing a highly automated construction pipeline that ensures real-time updates and generalization to a wider variety of terms. The dataset is updated annually and helps assess the performance of LLMs and potential improvement strategies.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hexuandeng/NewTerm.