This dataset contains cleaned, high-quality Norwegian Bokmål sentences containing numeric expressions harvested from both the Norwegian Colossal Corpus (NbAiLab/NCC) and the FineWeb-2 Norwegian Bokmål subset (HuggingFaceFW/fineweb-2 config nob_Latn).
This is a combined high-volume intermediate dataset built for the first stage of a template-based Norwegian synthetic speech (TTS) generation pipeline to improve number/digit… See the full description on the dataset page:
https://huggingface.co/datasets/pere/nb-asr-numerics-harvested.