Views
No views yet
<role>SYSTEM</role>{system_prompt}
detailed thinking on<|role_end|><role>HUMAN</role>{prompt}<|role_end|><role>ASSISTANT</role>
<think>| Filename | Quant type | File Size | Split | Description |
|---|---|---|---|---|
| Ling-3.0-tiny-bf16.gguf | bf16 | 15.80GB | false | Full BF16 weights. |
| Ling-3.0-tiny-Q8_0.gguf | Q8_0 | 8.41GB | false | Extremely high quality, generally unneeded but max available quant. |
| Ling-3.0-tiny-Q6_K_L.gguf | Q6_K_L | 6.96GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, recommended. |
| Ling-3.0-tiny-Q6_K.gguf | Q6_K | 6.84GB | false | Very high quality, near perfect, recommended. |
| Ling-3.0-tiny-Q5_K_L.gguf | Q5_K_L | 5.87GB | false | Uses Q8_0 for embed and output weights. High quality, recommended. |
| Ling-3.0-tiny-Q5_K_M.gguf | Q5_K_M | 5.72GB | false | High quality, recommended. |
| Ling-3.0-tiny-Q5_K_S.gguf | Q5_K_S | 5.55GB | false | High quality, recommended. |
| Ling-3.0-tiny-Q4_K_L.gguf | Q4_K_L | 5.10GB | false | Uses Q8_0 for embed and output weights. Good quality, recommended. |
| Ling-3.0-tiny-Q4_1.gguf | Q4_1 | 5.08GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| Ling-3.0-tiny-Q4_K_M.gguf | Q4_K_M | 4.92GB | false | Good quality, default size for most use cases, recommended. |
| Ling-3.0-tiny-Q4_K_S.gguf | Q4_K_S | 4.75GB | false | Slightly lower quality with more space savings, recommended. |
| Ling-3.0-tiny-Q4_0.gguf | Q4_0 | 4.62GB | false | Legacy format, kept for compatibility with older tools. |
| Ling-3.0-tiny-IQ4_NL.gguf | IQ4_NL | 4.62GB | false | Similar to IQ4_XS, but slightly larger. |
| Ling-3.0-tiny-IQ4_XS.gguf | IQ4_XS | 4.39GB | false | Decent quality, smaller than Q4_K_S with similar performance, recommended. |
| Ling-3.0-tiny-Q3_K_XL.gguf | Q3_K_XL | 4.13GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| Ling-3.0-tiny-IQ3_M.gguf | IQ3_M | 3.93GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| Ling-3.0-tiny-Q3_K_L.gguf | Q3_K_L | 3.91GB | false | Lower quality but usable, good for low RAM availability. |
| Ling-3.0-tiny-Q3_K_M.gguf | Q3_K_M | 3.79GB | false | Low quality. |
| Ling-3.0-tiny-IQ3_XS.gguf | IQ3_XS | 3.78GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| Ling-3.0-tiny-Q3_K_S.gguf | Q3_K_S | 3.64GB | false | Low quality, not recommended. |
| Ling-3.0-tiny-IQ3_XXS.gguf | IQ3_XXS | 3.46GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| Ling-3.0-tiny-Q2_K_L.gguf | Q2_K_L | 3.24GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| Ling-3.0-tiny-Q2_K.gguf | Q2_K | 3.00GB | false | Very low quality but surprisingly usable. |
| Ling-3.0-tiny-IQ2_M.gguf | IQ2_M | 2.83GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
hf download bartowski/Ling-3.0-tiny-GGUF --include "Ling-3.0-tiny-Q4_K_M.gguf" --local-dir ./pip install -U "huggingface_hub[cli]"hf download bartowski/Ling-3.0-tiny-GGUF --include "Ling-3.0-tiny-Q4_K_M.gguf" --local-dir ./curl -LsSf https://llama.app/install.sh | sh
llama-server -hf bartowski/Ling-3.0-tiny-GGUF:Q4_K_M--parse-special, so chat-format special tokens contribute to the importance matrix. The corpus rendered for this model is included in this repo: Ling-3.0-tiny-calibration-v6.txt. The imatrix is available here: Ling-3.0-tiny-imatrix.gguf.1{
2 "generator": "auto_quant_v2 calibration renderer",
3 "recipe": "calibration-v6",
4 "model": "Ling-3.0-tiny",
5 "encoder": "chat_template",
6 "chunk_size": 512,
7 "prose_chunks": 220,
8 "tool_chunks": 345,
9 "total_chunks": 565,
10 "tool_chunk_fraction": 0.611,
11 "n_conversations": 137,
12 "extension_convs_used": 0,
13 "conversation_token_lengths": [
14 523,
15 1594,
16 1193,
17 1476,
18 1046,
19 1300,
20 3127,
21 754,
22 1163,
23 1353,
24 1019,
25 2059,
26 836,
27 1200,
28 2755,
29 1189,
30 1099,
31 948,
32 694,
33 677,
34 1326,
35 990,
36 1308,
37 1167,
38 1839,
39 1463,
40 1601,
41 844,
42 1376,
43 1604,
44 1472,
45 1161,
46 1211,
47 1003,
48 1019,
49 1650,
50 1619,
51 1147,
52 433,
53 1912,
54 1392,
55 1048,
56 1355,
57 1973,
58 2023,
59 1230,
60 1569,
61 824,
62 2903,
63 1063,
64 2811,
65 723,
66 955,
67 915,
68 924,
69 655,
70 2396,
71 840,
72 1100,
73 1045,
74 1166,
75 1133,
76 868,
77 1151,
78 1114,
79 1530,
80 873,
81 1483,
82 2099,
83 803,
84 333,
85 1071,
86 3285,
87 2856,
88 671,
89 865,
90 974,
91 1022,
92 1244,
93 1052,
94 1074,
95 753,
96 1152,
97 983,
98 1244,
99 1468,
100 1321,
101 2041,
102 795,
103 608,
104 2714,
105 658,
106 1345,
107 1626,
108 1936,
109 1168,
110 581,
111 1336,
112 1136,
113 1653,
114 1759,
115 1625,
116 782,
117 961,
118 976,
119 2730,
120 697,
121 679,
122 709,
123 1354,
124 1011,
125 1544,
126 731,
127 361,
128 327,
129 2569,
130 947,
131 1085,
132 1815,
133 1970,
134 2651,
135 2644,
136 759,
137 931,
138 797,
139 884,
140 1190,
141 944,
142 809,
143 1266,
144 793,
145 668,
146 1711,
147 965,
148 880,
149 1240,
150 1409
151 ],
152 "warnings": []
153}