Views
No views yet
open-thoughts/OpenThinker3-7B quantized to FP8 (8-bit weights).Caveat. Needs compute capability >= 8.9 (Ada/Hopper+) to run fast.
| Source | open-thoughts/OpenThinker3-7B |
| Scheme | FP8 (8-bit) |
| Format | compressed-tensors |
| Parameters | 7.6B |
| Size on disk | 8.7 GB |
| Compression | 1.75x smaller than the 15.2 GB source |
| Left unquantized | lm_head |
| Quantized on | RTX 3090 |
| Quantized by | Sohailhosseini |
1vllm serve Sohailhosseini/OpenThinker3-7B-FP8 \
2 --max-model-len 327681from vllm import LLM, SamplingParams
2
3if __name__ == "__main__":
4 llm = LLM("Sohailhosseini/OpenThinker3-7B-FP8", max_model_len=32768)
5 out = llm.chat(
6 [{"role": "user", "content": "What is quantization? Answer in one sentence."}],
7 SamplingParams(temperature=0.6, max_tokens=512),
8 )
9 print(out[0].outputs[0].text)/v1/chat/completions endpoint - not merely loaded. Three prompts, greedy decoding, all three coherent.recipe.yaml in this repo is the exact modifier stack that was applied, and the scheme, ignored layers and hardware are in the table above.