Mechanistic robustness evaluation results for language models under six input
perturbations: character replacement, BPE-token replacement, word replacement,
local token shuffle, typographical corruption, and synonym replacement.
The repository is organized by model and perturbation:
models/
///evals.csv
The qwen2.5_1.5b/adversarial directory contains the separate adversarial
evaluation outputs and manifest. Failed or… See the full description on the dataset page:
https://huggingface.co/datasets/christian-hoang-04/decoding-robustness-results.