SpellBench is a benchmark for evaluating how well language models handle character-level and word-level linguistic operations. It includes 29,700 items across diverse tasks, with granular tracking of language and script for each sample.
Most LLMs operate at the token level and struggle with tasks that require reasoning about individual characters — spelling, reversing, counting letters, etc. SpellBench provides a… See the full description on the dataset page:
https://huggingface.co/datasets/omneity-labs/spellbench.