This is the ReazonSpeech corpus's large split, featuring Qwen3-ASR 1.7B transcriptions.
Since the original transcriptions often contain errors, comparing them with the Qwen3-ASR outputs could be useful.
Usage
import json
import webdataset as wds
from huggingface_hub import get_token