Views
No views yet
| Metric | Value |
|---|---|
| Base Model | openai/gpt-oss-20b |
| Architecture | Mixture-of-Experts Transformer |
| Total Parameters | ~20.3B (pruned from 21B) |
| Original Experts per Layer | 32 |
| Pruned Experts per Layer | 31 |
| Layers | 24 |
| Top-k Routing | 4 |
| Context Length | 128K tokens |
| Attention Heads | 64 (Query), 8 (Key-Value) |
| Residual Dimension | 2880 |
| Attention Pattern | Alternating dense & sliding window (128 tokens) |
| Positional Encoding | RoPE (Rotary Position Embedding) |
| Normalization | RMSNorm |
| Precision | BF16 |
| License | Apache 2.0 |
| Specialization | Instruction Following |
1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
3
4# Load the specialized model on CPU
5model = AutoModelForCausalLM.from_pretrained(
6 "AmanPriyanshu/gpt-oss-20.3b-specialized-instruction_following-pruned-moe-only-31-experts",
7 torch_dtype=torch.bfloat16,
8 device_map="cpu",
9 trust_remote_code=True
10)
11tokenizer = AutoTokenizer.from_pretrained("AmanPriyanshu/gpt-oss-20.3b-specialized-instruction_following-pruned-moe-only-31-experts")
12
13# Generate with the model
14messages = [
15 {"role": "user", "content": "Write a formal email to a professor requesting a meeting, including: subject line, greeting, purpose, proposed times, and professional closing."}
16]
17
18inputs = tokenizer.apply_chat_template(
19 messages,
20 add_generation_prompt=True,
21 return_tensors="pt",
22 return_dict=True,
23 reasoning_effort="medium"
24)
25
26# Ensure inputs are on the same device as model
27inputs = {k: v.to(model.device) for k, v in inputs.items()}
28
29outputs = model.generate(
30 **inputs,
31 max_new_tokens=512,
32 do_sample=True,
33 temperature=0.1,
34 top_p=0.9,
35 pad_token_id=tokenizer.eos_token_id,
36 eos_token_id=tokenizer.eos_token_id
37)
38
39# Decode only the generated part
40input_length = inputs['input_ids'].shape[1]
41response_tokens = outputs[0][input_length:]
42response = tokenizer.decode(response_tokens, skip_special_tokens=True)
43print(response)1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
3
4# Check MPS availability and load model
5device = "mps" if torch.backends.mps.is_available() else "cpu"
6
7model = AutoModelForCausalLM.from_pretrained(
8 "AmanPriyanshu/gpt-oss-20.3b-specialized-instruction_following-pruned-moe-only-31-experts",
9 torch_dtype=torch.float16, # Better MPS compatibility
10 device_map=device,
11 trust_remote_code=True,
12 low_cpu_mem_usage=True
13)
14tokenizer = AutoTokenizer.from_pretrained("AmanPriyanshu/gpt-oss-20.3b-specialized-instruction_following-pruned-moe-only-31-experts")
15
16# Generate with the model
17messages = [
18 {"role": "user", "content": "Write a formal email to a professor requesting a meeting, including: subject line, greeting, purpose, proposed times, and professional closing."}
19]
20
21inputs = tokenizer.apply_chat_template(
22 messages,
23 add_generation_prompt=True,
24 return_tensors="pt",
25 return_dict=True,
26 reasoning_effort="medium"
27)
28
29# Move inputs to model device
30inputs = {k: v.to(model.device) if isinstance(v, torch.Tensor) else v for k, v in inputs.items()}
31
32# Use torch.no_grad for MPS stability
33with torch.no_grad():
34 outputs = model.generate(
35 **inputs,
36 max_new_tokens=512,
37 do_sample=True,
38 temperature=0.1,
39 top_p=0.9,
40 pad_token_id=tokenizer.eos_token_id,
41 eos_token_id=tokenizer.eos_token_id,
42 use_cache=True
43 )
44
45# Decode only the generated part
46input_length = inputs['input_ids'].shape[1]
47response_tokens = outputs[0][input_length:]
48response = tokenizer.decode(response_tokens, skip_special_tokens=True)
49print(response)1device_map="auto" # Will automatically use GPU if available
2torch_dtype=torch.bfloat16 # or torch.float16<|channel|>analysis<|message|>Need formal email.<|end|><|start|>assistant<|channel|>final<|message|>**Subject:** Request for Meeting – Discussion of Research Project Progress
Dear Professor [Last Name],
I hope this message finds you well. I am writing to request a brief meeting with you to discuss the current status of my research project on [brief topic], as well as to seek your guidance on the next steps and potential publication avenues.
Could we meet at one of the following times next week?
- **Tuesday, Aug 20, 2025, 10:30 – 11:00 a.m.**
- **Wednesday, Aug 21, 2025, 2:00 – 2:30 p.m.**
- **Friday, Aug 23, 2025, 9:00 – 9:30 a.m.**
If none of these slots are convenient, I would be grateful if you could suggest an alternative time that fits your schedule.
Thank you very much for your time and consideration. I look forward to our conversation.
Sincerely,
[Your Full Name]
[Your Program / Department]
[University Name]
[Email Address] | [Phone Number]<|return|><|channel|>analysis<|message|>Need 6 steps numbered.<|end|><|start|>assistant<|channel|>final<|message|>1. **Gather ingredients and tools** – You’ll need eggs, a pinch of salt, a splash of milk or cream (optional), butter or oil, a bowl, a whisk or fork, and a non‑stick skillet with a spatula.
2. **Beat the eggs** – Crack the eggs into the bowl, add salt (and milk if using), then whisk or stir until the mixture is uniform and slightly frothy.
3. **Heat the skillet** – Place the skillet over medium‑low heat and add butter or oil, letting it melt and coat the surface without browning.
4. **Pour in the eggs** – Once the butter is foamy or the oil is shimmering, pour the beaten eggs into the skillet.
5. **Cook and stir gently** – Let the eggs sit for a few seconds until they start to set at the edges, then use the spatula to gently push the cooked portions toward the center while tilting the pan to let uncooked liquid flow to the edges. Repeat until the eggs are softly set but still slightly runny.
6. **Serve immediately** – Remove the skillet from heat, transfer the scrambled eggs to a plate, and enjoy while hot.<|return|><|channel|>analysis<|message|>Need 50 words exactly. Let's craft: "Water evaporates from oceans, lakes, and soil, forming clouds. Condensation cools vapor into droplets, which fall as precipitation—rain, snow, sleet, or hail—returning to Earth's surface. Runoff channels water into rivers, streams, and eventually back to oceans, while infiltration replenishes groundwater. This continuous cycle sustains life and regulates climate." Count words. Let's count: Water(1) evaporates2 from3 oceans,4 lakes,5 and6 soil,7 forming8 clouds.9 Condensation10 cools11 vapor12 into13 droplets,14 which15 fall16 as17 precipitation—rain,18 snow,19 sleet,20 or21 hail—returning22 to23 Earth's24 surface.25 Runoff26 channels27 water28 into29 rivers,30 streams,31 and32 eventually33 back34 to35 oceans,36 while37 infiltration38 replenishes39 groundwater.40 This41 continuous42 cycle43 sustains44 life45 and46 regulates47 climate48. Need 2 more words: "for"49 "all"50.<|end|><|start|>assistant<|channel|>final<|message|>Water evaporates from oceans, lakes, and soil, forming clouds. Condensation cools vapor into droplets, which fall as precipitation—rain, snow, sleet, or hail—returning to Earth's surface. Runoff channels water into rivers, streams, and eventually back to oceans, while infiltration replenishes groundwater. This continuous cycle sustains life and regulates climate for all.<|return|>1@misc{priyanshu2025gptoss,
2 title={{GPT-OSS MoE Expert Fingerprinting: Analyzing Expert Activation Patterns in Mixture of Experts Models}},
3 author={Priyanshu, Aman and Vijay, Supriti},
4 year={2025},
5 howpublished={\url{https://amanpriyanshu.github.io/GPT-OSS-MoE-ExpertFingerprinting/}},
6 note={Interactive analysis tool for expert activation patterns in MoE architectures}
7}