Views
No views yet
Revision note: originally converted from the FP8 release (meituan-longcat/LongCat-2.0-FP8), the current revision is re-converted from the bf16 master checkpoint.
| Tensor class | Bits | Parameters | Size | Share |
|---|---|---|---|---|
| Expert FFNs | 3 | 1.47 T | 598.5 GiB | 88.0% |
| N-gram embedding tables | 3 | 135 B | 55.0 GiB | 8.1% |
| Attention, dense MLPs | 6 | 31.4 B | 22.9 GiB | 3.4% |
Embeddings, lm_head | 6 | 2.8 B | 2.1 GiB | 0.3% |
| DSA indexer, MoE routers | 8 | 0.6 B | 0.6 GiB | 0.1% |
| Norms, correction biases (unquantized) | — | — | 0.9 GiB | 0.1% |
pip install git+https://github.com/ml-explore/mlx-lm.git@refs/pull/1464/head1from mlx_lm import load, generate
2
3model, tokenizer = load("kernelpool/LongCat-2.0-3bit-UVMAX")
4
5prompt = "hello"
6
7if tokenizer.chat_template is not None:
8 messages = [{"role": "user", "content": prompt}]
9 prompt = tokenizer.apply_chat_template(
10 messages, add_generation_prompt=True, return_dict=False,
11 )
12
13response = generate(model, tokenizer, prompt=prompt, verbose=True)