Finetuned following ALMA (
https://github.com/fe1ixxu/ALMA) on the Cantonese-Mandarin translation task.
Finetuning dataset: Sourced from the released raw dataset in
https://github.com/meganndare/cantonese-nlp
As the base model was already finetuned on Cantonese monolingual data, we only conducted finetuning on parallel sentences.
The ALMA code was linked as submodule.