Views
No views yet
| Metric | Value |
|---|---|
| Base Model | openai/gpt-oss-20b |
| Architecture | Mixture-of-Experts Transformer |
| Total Parameters | ~10.8B (pruned from 21B) |
| Original Experts per Layer | 32 |
| Pruned Experts per Layer | 15 |
| Layers | 24 |
| Top-k Routing | 4 |
| Context Length | 128K tokens |
| Attention Heads | 64 (Query), 8 (Key-Value) |
| Residual Dimension | 2880 |
| Attention Pattern | Alternating dense & sliding window (128 tokens) |
| Positional Encoding | RoPE (Rotary Position Embedding) |
| Normalization | RMSNorm |
| Precision | BF16 |
| License | Apache 2.0 |
| Specialization | Instruction Following |
1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
3
4# Load the specialized model on CPU
5model = AutoModelForCausalLM.from_pretrained(
6 "AmanPriyanshu/gpt-oss-10.8b-specialized-instruction_following-pruned-moe-only-15-experts",
7 torch_dtype=torch.bfloat16,
8 device_map="cpu",
9 trust_remote_code=True
10)
11tokenizer = AutoTokenizer.from_pretrained("AmanPriyanshu/gpt-oss-10.8b-specialized-instruction_following-pruned-moe-only-15-experts")
12
13# Generate with the model
14messages = [
15 {"role": "user", "content": "Write a formal email to a professor requesting a meeting, including: subject line, greeting, purpose, proposed times, and professional closing."}
16]
17
18inputs = tokenizer.apply_chat_template(
19 messages,
20 add_generation_prompt=True,
21 return_tensors="pt",
22 return_dict=True,
23 reasoning_effort="medium"
24)
25
26# Ensure inputs are on the same device as model
27inputs = {k: v.to(model.device) for k, v in inputs.items()}
28
29outputs = model.generate(
30 **inputs,
31 max_new_tokens=512,
32 do_sample=True,
33 temperature=0.1,
34 top_p=0.9,
35 pad_token_id=tokenizer.eos_token_id,
36 eos_token_id=tokenizer.eos_token_id
37)
38
39# Decode only the generated part
40input_length = inputs['input_ids'].shape[1]
41response_tokens = outputs[0][input_length:]
42response = tokenizer.decode(response_tokens, skip_special_tokens=True)
43print(response)1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
3
4# Check MPS availability and load model
5device = "mps" if torch.backends.mps.is_available() else "cpu"
6
7model = AutoModelForCausalLM.from_pretrained(
8 "AmanPriyanshu/gpt-oss-10.8b-specialized-instruction_following-pruned-moe-only-15-experts",
9 torch_dtype=torch.float16, # Better MPS compatibility
10 device_map=device,
11 trust_remote_code=True,
12 low_cpu_mem_usage=True
13)
14tokenizer = AutoTokenizer.from_pretrained("AmanPriyanshu/gpt-oss-10.8b-specialized-instruction_following-pruned-moe-only-15-experts")
15
16# Generate with the model
17messages = [
18 {"role": "user", "content": "Write a formal email to a professor requesting a meeting, including: subject line, greeting, purpose, proposed times, and professional closing."}
19]
20
21inputs = tokenizer.apply_chat_template(
22 messages,
23 add_generation_prompt=True,
24 return_tensors="pt",
25 return_dict=True,
26 reasoning_effort="medium"
27)
28
29# Move inputs to model device
30inputs = {k: v.to(model.device) if isinstance(v, torch.Tensor) else v for k, v in inputs.items()}
31
32# Use torch.no_grad for MPS stability
33with torch.no_grad():
34 outputs = model.generate(
35 **inputs,
36 max_new_tokens=512,
37 do_sample=True,
38 temperature=0.1,
39 top_p=0.9,
40 pad_token_id=tokenizer.eos_token_id,
41 eos_token_id=tokenizer.eos_token_id,
42 use_cache=True
43 )
44
45# Decode only the generated part
46input_length = inputs['input_ids'].shape[1]
47response_tokens = outputs[0][input_length:]
48response = tokenizer.decode(response_tokens, skip_special_tokens=True)
49print(response)1device_map="auto" # Will automatically use GPU if available
2torch_dtype=torch.bfloat16 # or torch.float16<|channel|>analysis<|message|>Need formal email.<|end|><|start|>assistant<|channel|>final<|message|>**Subject:** Request for Meeting – Discussion of Research Project Progress
Dear Professor [Last Name],
I hope this message finds you well. I would like to request a brief meeting to discuss the current status of my research project and to seek your guidance on the next steps.
Could we meet at one of the following times?
- Tuesday, April 12 at 10:00 a.m.
- Wednesday, April 13 at 2:00 p.m.
- Thursday, April 14 at 11:30 a.m.
If none of these slots work for you, I would be happy to adjust my schedule to accommodate your availability.
Thank you for your time and consideration. I look forward to our conversation.
Sincerely,
[Your Full Name]
[Program / Department]
[University]
[Phone] | [Email]<|return|><|channel|>analysis<|message|>Need 6 numbered steps.<|end|><|start|>assistant<|channel|>final<|message|>1. **Gather ingredients** – Get a bowl of eggs, a small amount of milk or cream, a handful of fresh or frozen vegetables, and a spoon or whisk.
2. **Heat the pan** – Turn a medium‑spot skillet or frying pan on to a gentle, steady heat on a stove or hot‑plate.
3. **Add the milk** – Pour the milk or cream into the pan, letting it spread evenly across the surface.
4. **Whate the eggs** – Crack the eggs into the pan and whisk them together with the milk, creating a smooth, slightly thickened mixture.
5. **Stheat the vegetables** – Sprinkle the vegetables over the mixture, letting them soften and mingle with the eggs.
6. **Serve** – Once the mixture is lightly browned and the vegetables are just beginning to melt, lift the pan, scoop the eggs into a bowl, and serve hot.<|return|><|channel|>analysis<|message|>Need 50 words exactly. Count. Let's craft: "Water originates in the atmosphere as precipitation, falls onto land, infiltrates soil, and moves through plants and organisms. It evaporates from surfaces, condensates into vapor, travels through the air, condensates again, and returns to the ground as rain. This continuous loop replenizes ecosystems, supports life, and sustens planetary balance." Count words. Let's count: Water(1) originates(2) in(3) the(4) atmosphere(5) as(6) precipitation,(7) falls(8) onto(9) land,(10) infiltrates(11) soil,(12) and(13) moves(14) through(15) plants(16) and(17) organisms.(18) It(19) evaporates(20) from(21) surfaces,(22) condensates(23) into(24) vapor,(25) travels(26) through(27) the(28) air,(29) condensates(30) again,(31) and(32) returns(33) to(34) the(35) ground(36) as(37) rain.(38) This(39) continuous(40) loop(41) replenizes(42) ecosystems,(43) supports(44) life,(45) and(46) sustens(47) planetary(48) balance.(49) Need one more word: "foreever" (50). Let's add "foreever" at end.<|end|><|start|>assistant<|channel|>final<|message|>Water originates in the atmosphere as precipitation, falls onto land, infiltrates soil, and moves through plants and organisms. It evaporates from surfaces, condensates into vapor, travels through the air, condensates again, and returns to the ground as rain. This continuous loop replenizes ecosystems, supports life, and sustens planetary balance foreever.<|return|>1@misc{priyanshu2025gptoss,
2 title={{GPT-OSS MoE Expert Fingerprinting: Analyzing Expert Activation Patterns in Mixture of Experts Models}},
3 author={Priyanshu, Aman and Vijay, Supriti},
4 year={2025},
5 howpublished={\url{https://amanpriyanshu.github.io/GPT-OSS-MoE-ExpertFingerprinting/}},
6 note={Interactive analysis tool for expert activation patterns in MoE architectures}
7}