Views
No views yet
| Metric | Value |
|---|---|
| Base Model | openai/gpt-oss-20b |
| Architecture | Mixture-of-Experts Transformer |
| Total Parameters | ~9.0B (pruned from 21B) |
| Original Experts per Layer | 32 |
| Pruned Experts per Layer | 12 |
| Layers | 24 |
| Top-k Routing | 4 |
| Context Length | 128K tokens |
| Attention Heads | 64 (Query), 8 (Key-Value) |
| Residual Dimension | 2880 |
| Attention Pattern | Alternating dense & sliding window (128 tokens) |
| Positional Encoding | RoPE (Rotary Position Embedding) |
| Normalization | RMSNorm |
| Precision | BF16 |
| License | Apache 2.0 |
| Specialization | Instruction Following |
1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
3
4# Load the specialized model on CPU
5model = AutoModelForCausalLM.from_pretrained(
6 "AmanPriyanshu/gpt-oss-9.0b-specialized-instruction_following-pruned-moe-only-12-experts",
7 torch_dtype=torch.bfloat16,
8 device_map="cpu",
9 trust_remote_code=True
10)
11tokenizer = AutoTokenizer.from_pretrained("AmanPriyanshu/gpt-oss-9.0b-specialized-instruction_following-pruned-moe-only-12-experts")
12
13# Generate with the model
14messages = [
15 {"role": "user", "content": "Write a formal email to a professor requesting a meeting, including: subject line, greeting, purpose, proposed times, and professional closing."}
16]
17
18inputs = tokenizer.apply_chat_template(
19 messages,
20 add_generation_prompt=True,
21 return_tensors="pt",
22 return_dict=True,
23 reasoning_effort="medium"
24)
25
26# Ensure inputs are on the same device as model
27inputs = {k: v.to(model.device) for k, v in inputs.items()}
28
29outputs = model.generate(
30 **inputs,
31 max_new_tokens=512,
32 do_sample=True,
33 temperature=0.1,
34 top_p=0.9,
35 pad_token_id=tokenizer.eos_token_id,
36 eos_token_id=tokenizer.eos_token_id
37)
38
39# Decode only the generated part
40input_length = inputs['input_ids'].shape[1]
41response_tokens = outputs[0][input_length:]
42response = tokenizer.decode(response_tokens, skip_special_tokens=True)
43print(response)1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
3
4# Check MPS availability and load model
5device = "mps" if torch.backends.mps.is_available() else "cpu"
6
7model = AutoModelForCausalLM.from_pretrained(
8 "AmanPriyanshu/gpt-oss-9.0b-specialized-instruction_following-pruned-moe-only-12-experts",
9 torch_dtype=torch.float16, # Better MPS compatibility
10 device_map=device,
11 trust_remote_code=True,
12 low_cpu_mem_usage=True
13)
14tokenizer = AutoTokenizer.from_pretrained("AmanPriyanshu/gpt-oss-9.0b-specialized-instruction_following-pruned-moe-only-12-experts")
15
16# Generate with the model
17messages = [
18 {"role": "user", "content": "Write a formal email to a professor requesting a meeting, including: subject line, greeting, purpose, proposed times, and professional closing."}
19]
20
21inputs = tokenizer.apply_chat_template(
22 messages,
23 add_generation_prompt=True,
24 return_tensors="pt",
25 return_dict=True,
26 reasoning_effort="medium"
27)
28
29# Move inputs to model device
30inputs = {k: v.to(model.device) if isinstance(v, torch.Tensor) else v for k, v in inputs.items()}
31
32# Use torch.no_grad for MPS stability
33with torch.no_grad():
34 outputs = model.generate(
35 **inputs,
36 max_new_tokens=512,
37 do_sample=True,
38 temperature=0.1,
39 top_p=0.9,
40 pad_token_id=tokenizer.eos_token_id,
41 eos_token_id=tokenizer.eos_token_id,
42 use_cache=True
43 )
44
45# Decode only the generated part
46input_length = inputs['input_ids'].shape[1]
47response_tokens = outputs[0][input_length:]
48response = tokenizer.decode(response_tokens, skip_special_tokens=True)
49print(response)1device_map="auto" # Will automatically use GPU if available
2torch_dtype=torch.bfloat16 # or torch.float16<|channel|>analysis<|message|>Need formal email.<|end|><|start|>assistant<|channel|>final<|message|>**Subject:** Request for Meeting – Discussion of Thesis Progress
Dear Professor [Last Name],
I hope this message finds you well. I would like to request a brief meeting to discuss the current status of my thesis and to seek your guidance on the next steps.
Could we meet on one of the following days?
- **Wednesday, 10 April – 10 :30 – 11 :00**
- **Friday, 15 April – 12 :30 – 12 :45**
- **Tuesday, 18 April – 15 :00 – 15 :15**
If any of these times are inconvenient, please let me know a suitable alternative.
Thank you for your continued support.
Sincerely,
[Your Full Name]
[Program / Department]
[University]
[Phone] | [Email]<|return|><|channel|>analysis<|message|>Need 6 numbered steps.<|end|><|start|>assistant<|channel|>final<|message|>1. **Gather the ingredients** – Grab a bowl of eggs, a small amount of butter or oil, and a handful of fresh or pre‑casted vegetables or seasonings.
2. **Heat the pan** – Turn the stove to medium‑high, place a skillet on the burner, and let it warm up to a gentle 100‑200 temperature.
3. **Add the butter or oil** – Spread the butter or oil across the pan, letting it coat the edges to keep the surface slick.
4. **Crate the eggs** – Crack each egg, let the yolies sit in the pan, and stir them together with a gentle whisk or a quick flick of the spoon.
5. **Fold the batter** – Let the eggs mingle with the butter, forming a soft, slightly golden swirl; keep the mixture from sticking by gently rotating the pan.
6. **Serve the scramble** – Flip the eggs into a spoonful, drizzle a pinch of salt or pepper, and plate the fluffy, savory breakfast for the hungry.<|return|><|channel|>analysis<|message|>Need 50 words exactly. Let's craft: "Water enters the atmosphere as vapor, condens into droplets, travels through the Earth's surface, evaporates, recondens, and returns to the atmosphere. This continuous cycle sustens life, shapes ecosystems, and maintains planetary balance." Count words. Let's count: Water(1) enters(2) the(3) atmosphere(4) as(5) vapor,(6) condens(7) into(8) droplets,(9) travels(10) through(11) the(12) Earth's(13) surface,(14) evaporates,(15) recondens,(16) and(17) returns(18) to(19) the(20) atmosphere.(21) This(22) continuous(23) cycle(24) sustens(24?) Wait "sustens" wrong. Use "s sustens"? Let's rewrite: "This continuous cycle sustens life, shapes ecosystems, and maintains planetary balance." Count: This(21) continuous(22) cycle(23) sustens(24) life,(25) shapes(26) ecosystems,(27) and(28) maintains(29) planetary(30) balance.(31) Total words: 31? Let's recount full: Water(1) enters(2) the(3) atmosphere(4) as(5) vapor,(6) condens(7) into(8) droplets,(9) travels(10) through(11) the(12) Earth's(13) surface,(14) evaporates,(15) recondens,(16) and(17) returns(18) to(19) the(20) atmosphere.(21) This(22) continuous(23) cycle(24) sustens(24?) Actually "s sustens" is wrong. Use "s sustens" not correct. Use "s sustens" maybe "s sustens" is wrong. Use "s sustens" no. Let's use "s sustens" no. Let's use "s sustens" no. Let's use "s sustens" no. Let's use "s sustens" no. Let's correct: "This continuous cycle sustens life, shapes ecosystems, and maintains planetary balance." Count: This(21) continuous(22) cycle(23) sustens(24) life,(25) shapes(26) ecosystems,(27) and(28) maintains(29) planetary(30) balance.(31) So total 311@misc{priyanshu2025gptoss,
2 title={{GPT-OSS MoE Expert Fingerprinting: Analyzing Expert Activation Patterns in Mixture of Experts Models}},
3 author={Priyanshu, Aman and Vijay, Supriti},
4 year={2025},
5 howpublished={\url{https://amanpriyanshu.github.io/GPT-OSS-MoE-ExpertFingerprinting/}},
6 note={Interactive analysis tool for expert activation patterns in MoE architectures}
7}