This is the BF16 version and cannot be hosted with vLLM. TensorRT-LLM is supported but not tested.
For the MXFP4 version that is vLLM compatible, check out
gpt-oss-120b-uncensored-mxfp4
Finetuning is done by LoRA on
Amazon FalseReject train set with 800 samples.
Evaluation results obtained on
Amazon FalseReject test set with 300 samples.
Code example, documentation, and further QAT checkpoints will be released soon.