Views
No views yet
| Model | #Total Params | #Activated Params | Context Length | Download |
|---|---|---|---|---|
| Ling-lite-base-1.5 | 16.8B | 2.75B | 128K | 🤗 HuggingFace |
| Ling-lite-1.5 | 16.8B | 2.75B | 128K | 🤗 HuggingFace |
| Benchmark | #shots | Ling-lite-1.5 | Ling-lite | Qwen3-4B-Instruct | Qwen3-8B-Instruct | Moonlight-16B-A3B-Instruct | LLaMA3.1-8B |
|---|---|---|---|---|---|---|---|
| MMLU(EM) | 5 | 74.33 | 71.27 | 70.09 | 75.97 | 70.74 | 68.67 |
| GPQA(Pass@1) | 0 | 36.55 | 29.73 | 40.4 | 47.10 | 19.51 | 27.59 |
| HumanEval(Pass@1) | 0 | 87.27 | 84.38 | 81.94 | 85.29 | 72.94 | 67.23 |
| LiveCodeBench 2408-2502 (Pass@1) | 0 | 22.7 | 18.94 | 21.8 | 26.88 | 14.76 | 18.41 |
| LCBench(pass@1) | 0 | 60.37 | 46.57 | 48.61 | 60.03 | 28.39 | 23.13 |
| Math(EM) | 0 | 82.62 | 72.80 | 81.46 | 82.70 | 67.1 | 52.42 |
| AIME2024(pass@1) | 0 | 21.88 | 10.21 | 20.62 | 26.25 | 6.88 | 7.29 |
| OlympiadBench(pass@1) | 0 | 52.30 | 36.44 | 54.33 | 56.11 | 32.85 | 17.04 |
| BBH(EM) | 0 | 75.75 | 66.38 | 78.21 | 79.33 | 63.45 | 68.05 |
| IFEval(Prompt Strict) | 0 | 77.70 | 77.99 | 81.06 | 83.55 | 49.01 | 73.01 |
| BFCL_live | 0 | 72.15 | 67.93 | 65.35 | 69.83 | 47.14 | 49.98 |

Needle In A Haystack (NIAH) tests. Ling-Lite-1.5 has improved long text generation capability and performs well across most context window lengths up to 128K.transformers:1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model_name = "inclusionAI/Ling-lite-1.5"
4
5model = AutoModelForCausalLM.from_pretrained(
6 model_name,
7 torch_dtype="auto",
8 device_map="auto"
9)
10tokenizer = AutoTokenizer.from_pretrained(model_name)
11
12prompt = "Give me a short introduction to large language models."
13messages = [
14 {"role": "system", "content": "You are Ling, an assistant created by inclusionAI"},
15 {"role": "user", "content": prompt}
16]
17text = tokenizer.apply_chat_template(
18 messages,
19 tokenize=False,
20 add_generation_prompt=True
21)
22model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
23
24generated_ids = model.generate(
25 **model_inputs,
26 max_new_tokens=512
27)
28generated_ids = [
29 output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
30]
31
32response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]@article{ling,
title = {Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs},
author = {Ling Team},
journal = {arXiv preprint arXiv:2503.05139},
year = {2025}
}