no post-training of any form has been performed on this model
postmortem
being honest: this model was not intended to be fully trained but sunk cost fallacy + curiousity made it so. loss is definitely better and was ~2x faster but seems less useful/same?