LaVIta is a parallel multitask SFT dataset covering three Italian varieties: Roman (Romanesco), Neapolitan (Napoletano), and Sicilian (Siciliano).
If you wish to extend this dataset with more languages, please have a look at the official repo.
Dataset Details
Dataset Description
LaVIta is the first SFT dataset targeting vernacular language understanding across three Italian varieties: Roman, Neapolitan, and Sicilian.
Rather… See the full description on the dataset page: https://huggingface.co/datasets/balthier7/LaVIta.