Views
No views yet
"chat_template_kwargs": {"enable_thinking": false} because my train data with none thinking. I also found low temperature ususually works well for code generation tasks.checkpoint-37000is the checkpoint where the model had just entered the plateau phase with a lower loss, and it might be better thancheckpoint-46000. However, the GGUF files I provide will still be released based on the final checkpoint at the end of training.
With OpenAI Compatible API (llama.cpp:llama-server)
1-> REQUEST ->
2
3{
4 "model": "XL-LuaCopilot-1.7B-FFT",
5 "messages": [
6 {"role": "system","content": "prefix"},
7 {"role": "user","content": "do\n--打印:你好世界\n local tex"},
8 {"role": "system","content": "suffix"},
9 {"role": "user","content": "nd"},
10 {"role": "system","content": "middle"}
11 ],
12 "stream": false,
13 "cache_prompt": false,
14 "samplers": "edkypmxt",
15 "temperature": 0.2,
16 "dynatemp_range": 0.1,
17 "dynatemp_exponent": 1,
18 "top_k": 20,
19 "top_p": 0.9,
20 "min_p": 0.05,
21 "typical_p": 1,
22 "xtc_probability": 0,
23 "xtc_threshold": 0.1,
24 "repeat_last_n": 32,
25 "repeat_penalty": 1.1,
26 "presence_penalty": 0,
27 "frequency_penalty": 0.5,
28 "dry_multiplier": 0,
29 "dry_base": 1.75,
30 "dry_allowed_length": 2,
31 "dry_penalty_last_n": -1,
32 "max_tokens": -1,
33 "timings_per_token": true,
34 "chat_template_kwargs": {"enable_thinking": false}
35}
36
37-> RESPONSE ->
38
39{
40 "choices": [
41 {
42 "finish_reason": "stop",
43 "index": 0,
44 "message": {
45 "role": "assistant",
46 "content": "<think>\n\n</think>\n\nt = \"你好世界\"\n print(text)\ne"
47 }
48 }
49 ],
50 ...
51}I know Qwen has<|fim_prefix|>/<|fim_suffix|>/<|fim_middle|>tokens, but I'm not sure Qwen3 trains these tokens (I just know Qwen2.5-Coder does). To use code generation easily, I use chatml format.
If you just want to chat with it, you can use some tricks like this:
<|im_end|>
<|im_start|>system
prefix<|im_end|>
<|im_start|>user
do
--打印:你好世界
local tex<|im_end|>
<|im_start|>system
suffix<|im_end|>
<|im_start|>user
nd<|im_end|>
<|im_start|>system
middle<|im_start|>user
<|im_end|>
<|im_start|>system
prefix<|im_end|>
<|im_start|>user
do
--打印:你好世界
local tex<|im_end|>
<|im_start|>system
suffix<|im_end|>
<|im_start|>user
nd<|im_end|>
<|im_start|>system
middle<|im_end|><|im_start|>user\n<|im_end|> part.Online GPU is Expensive !
| 类别 | 配置详情 |
|---|---|
| 镜像 | Ubuntu 22.04 |
| PyTorch | 2.5.1 |
| Python | 3.12 |
| CUDA | 12.4 |
| GPU | RTX 4090 (24GB) * 1 |
| CPU | 25 vCPU Intel(R) Xeon(R) Platinum 8481C |
| 内存 | 90GB |
| 硬盘 | 30 GB + 50 GB |
| 时长 | 3 Day |