Views
No views yet
ornith-ai/Ornith-1.0-9B quantized to FP8 (8-bit weights).Caveat. Needs compute capability >= 8.9 (Ada/Hopper+) to run fast.
| Source | ornith-ai/Ornith-1.0-9B |
| Scheme | FP8 (8-bit) |
| Format | compressed-tensors |
| Parameters | 9.4B |
| Size on disk | 11.9 GB |
| Compression | 1.58x smaller than the 18.8 GB source |
| Left unquantized | lm_head, re:.*visual.*, re:.*vision_tower.*, re:.*vision_model.*, re:.*vision.*, re:.*multi_modal_projector.*, re:.*merger.* |
| Quantized on | L40S |
| Quantized by | Sohailhosseini |
1vllm serve Sohailhosseini/Ornith-1.0-9B-FP8 \
2 --max-model-len 327681from vllm import LLM, SamplingParams
2
3if __name__ == "__main__":
4 llm = LLM("Sohailhosseini/Ornith-1.0-9B-FP8", max_model_len=32768)
5 out = llm.chat(
6 [{"role": "user", "content": "What is quantization? Answer in one sentence."}],
7 SamplingParams(temperature=0.6, max_tokens=512),
8 )
9 print(out[0].outputs[0].text)recipe.yaml in this repo is the exact modifier stack that was applied, and the scheme, ignored layers and hardware are in the table above.