This dataset contains synthetic Portuguese (Brazil) speech/text examples for an experimental omni speech pipeline:
Whisper/ASR -> Tucano hidden states -> Qwen3 audio-code talker -> Qwen3-TTS tokenizer decoder
fast_pass_answer_texts.jsonl: 50,000 synthetic Q&A/text rows.
teacher_wavs_det_v1/: 12,204 generated WAV files.
audio_metadata.jsonl: lightweight audio manifest with relative WAV paths and text.… See the full description on the dataset page:
https://huggingface.co/datasets/marcosremar2/qwen3-omni-ptbr-qwen3tts-synthetic.