CrossER is a benchmark for context-dependent cross-system entity resolution where surface features are deliberately misleading. Match pairs average only 0.29 string similarity (names look unrelated), while non-match pairs average 0.94 similarity (names look identical).
In real enterprises, matching Product 4418 to Maltodextrin DE20 Grade A requires consulting migration runbooks, classification guides… See the full description on the dataset page:
https://huggingface.co/datasets/smurthy5/CrossER.