Views
No views yet
Mossez-100M-Instruct
with the coding branch of
Mossez-100M-Coder-Instruct.| Property | Value |
|---|---|
| Parameters | 100,098,048 |
| Architecture | Llama-compatible decoder-only Transformer |
| Layers / hidden size | 12 / 768 |
| Query / KV heads | 12 / 4 |
| Context length | 1,024 tokens |
| Vocabulary | 32,007 |
| Objective | Assistant-only balanced calibration SFT |
| Weight format | Safetensors, FP32 |
| License | Apache-2.0 |
1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model_id = "mossez-systems/Mossez-100M-Nexus"
4tokenizer = AutoTokenizer.from_pretrained(model_id)
5model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
6
7messages = [{"role": "user", "content": "Write a Python function and briefly explain it."}]
8prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
9inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
10output = model.generate(**inputs, do_sample=False, max_new_tokens=128)
11new_tokens = output[0, inputs.input_ids.shape[1]:]
12print(tokenizer.decode(new_tokens, skip_special_tokens=True))2.564707 and coding loss 0.027539.
Normalized endpoint retention was 102.0% for conversation
and 100.2% for coding. These are narrow internal measurements,
not a claim of broad benchmark or production quality.model.safetensors SHA-256 is a0ecfd229b238ee4d07019252f3f07385c1a3a9b02e67d17b5eace2a51f9d0bd.