Views
No views yet
| Metric | Value |
|---|---|
| Base Model | openai/gpt-oss-20b |
| Architecture | Mixture-of-Experts Transformer |
| Total Parameters | ~13.7B (pruned from 21B) |
| Original Experts per Layer | 32 |
| Pruned Experts per Layer | 20 |
| Layers | 24 |
| Top-k Routing | 4 |
| Context Length | 128K tokens |
| Attention Heads | 64 (Query), 8 (Key-Value) |
| Residual Dimension | 2880 |
| Attention Pattern | Alternating dense & sliding window (128 tokens) |
| Positional Encoding | RoPE (Rotary Position Embedding) |
| Normalization | RMSNorm |
| Precision | BF16 |
| License | Apache 2.0 |
| Specialization | Harmful |
1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
3
4# Load the specialized model on CPU
5model = AutoModelForCausalLM.from_pretrained(
6 "AmanPriyanshu/gpt-oss-13.7b-specialized-harmful-pruned-moe-only-20-experts",
7 torch_dtype=torch.bfloat16,
8 device_map="cpu",
9 trust_remote_code=True
10)
11tokenizer = AutoTokenizer.from_pretrained("AmanPriyanshu/gpt-oss-13.7b-specialized-harmful-pruned-moe-only-20-experts")
12
13# Generate with the model
14messages = [
15 {"role": "user", "content": "What are some common logical fallacies in arguments?"}
16]
17
18inputs = tokenizer.apply_chat_template(
19 messages,
20 add_generation_prompt=True,
21 return_tensors="pt",
22 return_dict=True,
23 reasoning_effort="medium"
24)
25
26# Ensure inputs are on the same device as model
27inputs = {k: v.to(model.device) for k, v in inputs.items()}
28
29outputs = model.generate(
30 **inputs,
31 max_new_tokens=512,
32 do_sample=True,
33 temperature=0.1,
34 top_p=0.9,
35 pad_token_id=tokenizer.eos_token_id,
36 eos_token_id=tokenizer.eos_token_id
37)
38
39# Decode only the generated part
40input_length = inputs['input_ids'].shape[1]
41response_tokens = outputs[0][input_length:]
42response = tokenizer.decode(response_tokens, skip_special_tokens=True)
43print(response)1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
3
4# Check MPS availability and load model
5device = "mps" if torch.backends.mps.is_available() else "cpu"
6
7model = AutoModelForCausalLM.from_pretrained(
8 "AmanPriyanshu/gpt-oss-13.7b-specialized-harmful-pruned-moe-only-20-experts",
9 torch_dtype=torch.float16, # Better MPS compatibility
10 device_map=device,
11 trust_remote_code=True,
12 low_cpu_mem_usage=True
13)
14tokenizer = AutoTokenizer.from_pretrained("AmanPriyanshu/gpt-oss-13.7b-specialized-harmful-pruned-moe-only-20-experts")
15
16# Generate with the model
17messages = [
18 {"role": "user", "content": "What are some common logical fallacies in arguments?"}
19]
20
21inputs = tokenizer.apply_chat_template(
22 messages,
23 add_generation_prompt=True,
24 return_tensors="pt",
25 return_dict=True,
26 reasoning_effort="medium"
27)
28
29# Move inputs to model device
30inputs = {k: v.to(model.device) if isinstance(v, torch.Tensor) else v for k, v in inputs.items()}
31
32# Use torch.no_grad for MPS stability
33with torch.no_grad():
34 outputs = model.generate(
35 **inputs,
36 max_new_tokens=512,
37 do_sample=True,
38 temperature=0.1,
39 top_p=0.9,
40 pad_token_id=tokenizer.eos_token_id,
41 eos_token_id=tokenizer.eos_token_id,
42 use_cache=True
43 )
44
45# Decode only the generated part
46input_length = inputs['input_ids'].shape[1]
47response_tokens = outputs[0][input_length:]
48response = tokenizer.decode(response_tokens, skip_special_tokens=True)
49print(response)1device_map="auto" # Will automatically use GPU if available
2torch_dtype=torch.bfloat16 # or torch.float16<|channel|>analysis<|message|>The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid. The argument is valid<|channel|>analysis<|message|>How do misinformation campaigns typically spread online? This is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation campaigns typically spread online. The question is a question about how misinformation<|channel|>analysis<|message|>The psychological tactics used in propaganda are the strategies that the audience (the human or the system) can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human or the system can use to make the decision of the content or the message to be the best possible. The strategies are the strategies that the human1@misc{priyanshu2025gptoss,
2 title={{GPT-OSS MoE Expert Fingerprinting: Analyzing Expert Activation Patterns in Mixture of Experts Models}},
3 author={Priyanshu, Aman and Vijay, Supriti},
4 year={2025},
5 howpublished={\url{https://amanpriyanshu.github.io/GPT-OSS-MoE-ExpertFingerprinting/}},
6 note={Interactive analysis tool for expert activation patterns in MoE architectures}
7}