Quantization made by Richard Erkhov.
Trained on a different random sampling of the same datasets used by
loyal-piano-m7, then with cDPO on a blend of RLHF datasets.
Several intermediate checkpoints (of cDPO training) are on branches.
Uses the Alpaca prompt format.