This artifact pairs a GPTQ target model with an EAGLE3 draft model for speculative decoding.
Performance
The following measurements characterize text generation on Modalix.
Measured with MoLE using batch size 1, five samples per input length, and up to 128 generated tokens. Values are arithmetic means. TTFT includes language-model prefill and the first generated token; generation rate is measured after the first token.
Input tokens
Mean TTFT (seconds)
Mean generation rate (tokens/second)
128
0.24
18.90
256
0.47
18.38
512
0.93
15.22
1024
1.90
16.83
2048
4.12
14.69
3072
7.23
12.91
4096
10.77
11.27
5120
15.42
7.49
6144
20.59
5.92
7168
27.51
4.64
Prerequisites
To run this model, you need a SiMa.ai Modalix device with the SiMa.ai Neat Runtime installed.
Installation and Deployment
Download the precompiled model directly on Modalix: