100M-token slice of HuggingFaceFW/fineweb-2 (tam_Taml) for LAS tokenizer training under the MultilingualUnigramLM org.
Source: HuggingFaceFW/fineweb-2
Config: tam_Taml
Documents: 263,523
Whitespace-tokens: 100,000,109
Chars: 894,087,067
Format: single data/train-00000-of-00001.parquet, schema (text: string). Mirrors the LangMap-TheStack-*-100M layout.