Modelscope without the watermark, trained in 320x320 from the
original weights, with no skipped frames for less flicker.
See comparison here:
https://www.youtube.com/watch?v=r4tOc30Zu0w
Model was trained on a subset of the vimeo90k dataset + a selection of music videos