French conversational-speech dataset used to fine-tune
FrWhisper, a Whisper Large-V3 model
adapted to spoken French, including interjections, hesitations, and other
natural speech phenomena. It combines material from two French corpora,
ESLO and
LangAge, segmented into
utterance-level clips paired with their transcriptions.
This release provides transcripts + audio only (16 kHz mono); it does not
include precomputed log-mel features.… See the full description on the dataset page:
https://huggingface.co/datasets/aihpi/FrWhisper-dataset.