Data Pipeline:
https://github.com/PoTaTo-Mika/Shore-Data-Engine
对于早期的小型专辑,都是wav格式方便直接使用。对于大规模(10000+ hours)的系列,都是opus进行保存,请注意储存空间。
请考虑在使用该数据集的论文中引用我们的工作:
@misc{cheng2025mikupalautomatedstandardizedmultimodal,
title={MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling},
author={Yifan Cheng and… See the full description on the dataset page:
https://huggingface.co/datasets/PoTaTo721/Shore-Lunch-Box.