CAESAR-TINY is a synthetic code-switched dataset generated by combining monolingual samples in Catalan and Spanish.
The process includes trimming silences, normalizing audio volume, and introducing random pauses. It contains 2 hours of speech data, created by concatenating audio from the Common voice 17 Benchmark split and VoxForge Spanish datasets.