Dataset Preprocessing for 10M and 100M Text-Only Tracks
Overview
This document describes the preprocessing steps applied to the datasets used for the 10M and 100M text-only tracks. The datasets are a mixture of 10 different corpora, as shown in Table 1 below.