Multimodal piano dataset of 6,019 short excerpts (3–8 measures each) cut from the
ASAP piano dataset. Each row aligns four
modalities for the same musical window:
column
type
description
audio
Audio
excerpt audio, 22 050 Hz mono FLAC
kern
string
full **kern score text of the excerpt (**kern header included)
st_plus
Sequence[str]
ScoreTransformer+ token list (bar/key/time/note/chord/rest/clef)