RESD with transcripts and separate text-emotion labels.
RESD was recorded in a studio by 20 voice actors. There was no script: the actors were not handed lines to read. Instead each actor in a pair was privately given an emotion to play, and the dialogue was improvised from there. So the words are spontaneous while the emotion is deliberate — which is the point, and also the limit. The label describes what the actor was… See the full description on the dataset page:
https://huggingface.co/datasets/Aniemore/resd_annotated_multi.