Derived data + per-seed evaluation results for the
Beyond Transcript Alignment research project (frozen-frozen
speech-to-LLM adapters under counterfactual training on the
StressTest benchmark).
Code:
https://github.com/Nurgali-Kadyrbek/frozen-speech-llm-stress
cf_pairs/cf_pairs_train.jsonl — 3666 same-transcript counterfactual
pairs from Stress-17K-raw probe-train (transcript IDs + stress
indices; no raw audio).… See the full description on the dataset page:
https://huggingface.co/datasets/nur-dev/stress17k-counterfactual-pairs.