Views
No views yet

Qwen3_5ForConditionalGeneration (linear + full attention interleaved, Gated Delta Net; text path extracted for GGUF)| File | Bits | Notes |
|---|---|---|
Ornstein-3.6-27B-Q8_0.gguf | 8 | Reference, near-lossless |
Ornstein-3.6-27B-Q6_K.gguf | 6.5 | Great default for 32 GB+ systems |
Ornstein-3.6-27B-Q5_K_M.gguf | 5.5 | Excellent quality/size balance |
Ornstein-3.6-27B-Q5_K_S.gguf | 5.5 | Slightly smaller Q5 |
Ornstein-3.6-27B-Q5_0.gguf | 5 | Legacy 5-bit |
Ornstein-3.6-27B-Q4_K_M.gguf | 4.5 | Common 24 GB-card default |
Ornstein-3.6-27B-Q4_K_S.gguf | 4.5 | Smaller Q4 |
Ornstein-3.6-27B-Q4_0.gguf | 4 | Legacy 4-bit |
Ornstein-3.6-27B-IQ4_NL.gguf | 4.25 | Non-linear 4-bit I-quant |
Ornstein-3.6-27B-IQ4_XS.gguf | 4.25 | Smaller than Q4_K_S, comparable quality |
Ornstein-3.6-27B-Q3_K_L.gguf | 3.5 | Largest Q3 |
Ornstein-3.6-27B-Q3_K_M.gguf | 3.5 | Usable; quality below Q4 |
Ornstein-3.6-27B-Q3_K_S.gguf | 3.5 | Smaller Q3 |
Ornstein-3.6-27B-IQ3_M.gguf | 3.3 | Mixed I-quant, beats Q3_K_S at similar size |
Ornstein-3.6-27B-IQ3_S.gguf | 3.1 | 3-bit I-quant |
Ornstein-3.6-27B-IQ3_XS.gguf | 3.0 | Smaller 3-bit I-quant |
Ornstein-3.6-27B-IQ3_XXS.gguf | 3.0 | Aggressive 3-bit |
Ornstein-3.6-27B-Q2_K.gguf | 2.6 | Lowest K-quant; expect degraded quality |
1# Interactive chat
2llama-cli -m Ornstein-3.6-27B-Q4_K_M.gguf -cnv
3
4# Single prompt
5llama-cli -m Ornstein-3.6-27B-Q5_K_M.gguf -p "Write a haiku about hybrid attention."
6
7# OpenAI-compatible server
8llama-server -m Ornstein-3.6-27B-Q4_K_M.gguf --host 0.0.0.0 --port 8080 -c 8192Modelfile), koboldcpp, and text-generation-webui all load these GGUFs provided their bundled llama.cpp supports Qwen3_5ForConditionalGeneration with Gated Delta Net.1# 1. Convert safetensors → BF16 GGUF
2python llama.cpp/convert_hf_to_gguf.py <model_dir> \
3 --outtype bf16 --outfile Ornstein-3.6-27B-BF16.gguf
4
5# 2. Quantize (example)
6llama-quantize Ornstein-3.6-27B-BF16.gguf \
7 Ornstein-3.6-27B-Q4_K_M.gguf Q4_K_M