Mixtral-8x22B with a bit of tensor surgery/lobotomy to chop out all but one of the MoE FFNs and configured to act like a dense model. Is it useful? Not in this state, without further training of the FFN or merging a few models together. When testing it, the only tokens it wanted to output was space and carriage return. Enjoy.