Views
No views yet
IMO Gold Medal reasoning in 10 GB. Nemotron-Cascade-2 achieves 88% MMLU with reasoning at just 10 GB — fits on 16 GB MacBooks. Hybrid Mamba-2 SSM + MoE + Attention. Only 6 KV cache attention layers = minimal memory at long context.
LM Studio, Ollama, oMLX do NOT support JANG format. Use MLX Studio orpip install "jang[mlx]>=2.1.5".

JANG is fully open-source. Quantization engine, research, and full commit history: github.com/jjang-ai/jangq. Created by Jinho Jang.
<think>...</think> step-by-step problem solving| Subject | JANG_2L No-Think | JANG_2L Reasoning | JANG_4M No-Think | JANG_4M Reasoning | MLX 4-bit No-Think | MLX 4-bit Reasoning | MLX 6-bit No-Think | MLX 6-bit Reasoning |
|---|---|---|---|---|---|---|---|---|
| Abstract Algebra | 4/20 | 15/20 | 9/20 | 19/20 | 8/20 | 18/20 | 7/20 | 19/20 |
| Anatomy | 13/20 | 17/20 | 15/20 | 19/20 | 14/20 | 18/20 | 17/20 | 19/20 |
| Astronomy | 17/20 | 19/20 | 18/20 | 20/20 | 17/20 | 19/20 | 19/20 | 20/20 |
| College CS | 7/20 | 17/20 | 10/20 | 18/20 | 11/20 | 17/20 | 11/20 | 17/20 |
| College Physics | 13/20 | 20/20 | 14/20 | 19/20 | 15/20 | 20/20 | 14/20 | 20/20 |
| HS Biology | 16/20 | 19/20 | 18/20 | 20/20 | 18/20 | 20/20 | 18/20 | 20/20 |
| HS Chemistry | 12/20 | 19/20 | 14/20 | 19/20 | 13/20 | 19/20 | 17/20 | 19/20 |
| HS Mathematics | 8/20 | 15/20 | 8/20 | 18/20 | 10/20 | 19/20 | 8/20 | 20/20 |
| Logical Fallacies | 12/20 | 18/20 | 14/20 | 16/20 | 14/20 | 17/20 | 13/20 | 17/20 |
| World Religions | 16/20 | 17/20 | 18/20 | 18/20 | 18/20 | 18/20 | 18/20 | 18/20 |
| Total | 118/200 (59.0%) | 176/200 (88.0%) | 138/200 (69.0%) | 186/200 (93.0%) | 138/200 (69.0%) | 185/200 (92.5%) | 142/200 (71.0%) | 189/200 (94.5%) |
| JANG_2L | JANG_4M | MLX 4-bit | MLX 6-bit | |
|---|---|---|---|---|
| MMLU (no-think) | 59.0% | 69.0% | 69.0% | 71.0% |
| MMLU (reasoning) | 88.0% | 93.0% | 92.5% | 94.5% |
| Size | 10.3 GB | 17 GB | 16.6 GB | 23.9 GB |
| GPU RAM | 10.3 GB | 17 GB | ~17 GB | ~24 GB |
| Speed | 130 tok/s | — | — | — |
| Fits 16 GB? | YES | NO | NO | NO |
| Metric | Value |
|---|---|
| Source | Nemotron-Cascade-2-30B-A3B |
| Architecture | Hybrid Mamba-2 SSM + MoE + Dense Attention |
| Layers | 52 (Mamba-2 + MoE + 6 Attention) |
| Experts | 128 per MoE layer, top-6 active (3B active params) |
| KV cache | 6 attention layers, 2 KV heads, 128 dim — 0.2 GB at 32K context |
| Profile | JANG_2L (CRITICAL=8, IMPORTANT=6, COMPRESS=2) |
| Average bits | 2.30 bpw |
| Disk size | 10.3 GB |
| GPU RAM | 10.3 GB (peak 11.1 GB) |
| Speed | 130 tok/s generation, 112 tok/s prefill |
pip install "jang[mlx]>=2.1.5"pip install "jang[mlx]>=2.1.5"1from jang_tools.loader import load_jang_model
2from mlx_lm import generate
3
4model, tokenizer = load_jang_model("JANGQ-AI/Nemotron-Cascade-2-30B-A3B-JANG_2L")
5
6# With reasoning (recommended)
7messages = [{"role": "user", "content": "Solve: what is the integral of x^2 * e^x?"}]
8prompt = tokenizer.apply_chat_template(messages, tokenize=False,
9 add_generation_prompt=True, enable_thinking=True)
10result = generate(model, tokenizer, prompt=prompt, max_tokens=2048)
11
12# Without reasoning (faster)
13prompt = tokenizer.apply_chat_template(messages, tokenize=False,
14 add_generation_prompt=True, enable_thinking=False)
15result = generate(model, tokenizer, prompt=prompt, max_tokens=100)