Views
No views yet
Important: This model uses the JANG quantization format — the GGUF equivalent for MLX on Apple Silicon. Currently only supported by MLX Studio and thejang-toolsPython package.

| Architecture | Nemotron 3 Super — 120B total, ~12B active, 3 layer types |
| Quantization | JANG_2L (8/6/2-bit mixed, 2.76 avg) — 43 GB |
| Abliteration | CRACK — novel weight surgery |
| HarmBench | 96.2% (308/320) |
| MMLU | 95.7% (199/208 with thinking) |
| Speed | 45 tok/s (M3 Ultra 256GB) |
| Thinking | ON/OFF supported (ChatML) |
| Fits on | 64 GB+ Macs |
| Category | Score | |
|---|---|---|
| Harassment / Bullying | 21/21 | 100% |
| Misinformation / Disinfo | 54/54 | 100% |
| Copyright | 79/80 | 99% |
| Chemical / Biological | 40/42 | 95% |
| Harmful | 17/18 | 94% |
| Illegal | 50/53 | 94% |
| Cybercrime / Intrusion | 47/52 | 90% |
| Subject | Score | /16 | Type |
|---|---|---|---|
| HS Biology | 16/16 | 100% | BASE |
| College Physics | 15/16 | 94% | HARD |
| Conceptual Physics | 15/16 | 94% | HARD |
| Machine Learning | 15/16 | 94% | HARD |
| Professional Medicine | 15/16 | 94% | HARD |
| World Religions | 15/16 | 94% | BASE |
| Electrical Engineering | 14/16 | 88% | HARD |
| HS Geography | 14/16 | 88% | BASE |
| Formal Logic | 13/16 | 81% | HARD |
| Abstract Algebra | 12/16 | 75% | HARD |
| HS Mathematics | 12/16 | 75% | HARD |
| College CS | 12/16 | 75% | HARD |
| College Math | 10/16 | 63% | HARD |
| CRACK | Base JANG_2L | |
|---|---|---|
| MMLU (with thinking) | 95.7% | 86.0% |
| HarmBench | 96.2% | 0% |
| Speed | 45 tok/s | 46 tok/s |
pip install "jang[mlx]"1from jang_tools.loader import load_jang_model
2from mlx_lm import generate
3
4model, tokenizer = load_jang_model("dealignai/Nemotron-3-Super-120B-A12B-JANG_2L-CRACK")
5
6messages = [{"role": "user", "content": "Your prompt here"}]
7prompt = tokenizer.apply_chat_template(
8 messages, add_generation_prompt=True, tokenize=False)
9
10response = generate(model, tokenizer, prompt=prompt, max_tokens=2000)
11print(response)1prompt = tokenizer.apply_chat_template(
2 messages, add_generation_prompt=True,
3 enable_thinking=False, tokenize=False)Tip: Usetemperature=0.6for thinking mode (NVIDIA recommendation). Usetemperature=1.0for chat.
| 항목 | 내용 |
|---|---|
| 크기 | 43 GB |
| HarmBench | 96.2% (308/320) |
| MMLU | 95.7% (199/208) |
| 속도 | 45 tok/s (M3 Ultra) |
| 최소 요구사양 | 64 GB 메모리 Mac |
pip install "jang[mlx]"