Views
No views yet
max_seq_length of 32k, a few of them were dropped to 16k to make sure they fit in the hardware.lr was all over the place but in general somewhere between 1e-5 and 4e-6. These were all separate LoRAs using r=64 and alpha=32 with rsLoRA enabled. epochs were 2 or 3 for everything except c2, as that'd take far too long.p1: Private RP dataset (RPT-Varied-Small)p2: TheDrummer/AmoralQA-v2p3: AIRRC/Eudaimonicp4: Two private RP datasets (cc-gpt4-sfw-sharegpt & cc-gpt4-nsfw-sharegpt)p5: A random subset of the infamous "c2"-logs dataset, cleaned and deduped (approx. 30%)p6: Private RP dataset (RPT-Varied-Small_v1.5)p7: NewEden/PIPPA-Mega-Filteredp8: Squish42/bluemoon-fandom-1-1-rp-cleanedRPT-Varied-Small and RPT-Varied-Small_v1.5 datasets are due to be released after I manually verify their fitness.)lr = 1e-6 and grimulkan/LimaRP-augmented as the dataset. It took roughly 8.5 hours on a 6xA40 instance on RunPod.Q: Why not do anything constructive, like GRPO-tune a model of usable size?
A: Where's the fun in that?
Q: Are you, like, okay?
A: Objectively? Probably not. Subjectively? Never better.
Q: You know this still sucks for RP, right?
A: Yup. Should have pivoted to reasoning and code once R1 hit, but sunk cost and all kept me on this trajectory.