Views
No views yet
Qwen3-30B-A3B-Thinking-2507 (Base Model, 4-bit)Qwen3-30B-FT-LoRA (Our fine-tuned adapter)"list bubble function sort python a to Write"
see_layers_30B_T.py script, which analyzes the model's prediction for a specific token (we chose the 20th token) across all 48 layers of the network.Generated Snippet:Okay, the user is asking for a list bubble function sort in Python. Let me check if I understand the query correctly.
Generated Snippet:We are going to write a bubble sort function in Python that sorts a list from A to Z (ascending order).
understand was already the #2 candidate. The model had formed a hypothesis early on.7.12 to 0.40. The probability of understand skyrocketed to 87.56%. The decision was made, cleanly and decisively.L-20 | gMaps (9%) | #3 | 6.76% | 7.12 | ·understanding ·understand BeginInit ·correctly
...
L-44 | ·understand (88%) | ✅ #1 | 87.56% | 0.40 | ·understood ·understands 理解 ·understandingascending was ranked a miserable #7, with only 1.26% probability. The model was lost.ascending finally crawled to the #1 spot. The decision was a last-second guess, not a confident conclusion.L-40 | 也就是 (57%) | #7 | 1.26% | 2.29 | ·ascending ·alphabetical つまり 即
...
L-47 | ·ascending (50%) | ✅ #1 | 49.54% | 0.76 | ascending Ascending i ·Asc1python see_layers_30B_T.py
2🦥 Unsloth: Will patch your computer to enable 2x faster free finetuning.
3🦥 Unsloth Zoo will now patch everything to make training faster!
4🚀 Loading LoRA Model: ./tmodels/Qwen3-30B-A3B-Thinking-2507-4bit-FT-lora
5==((====))== Unsloth 2026.1.4: Fast Qwen3_MoE patching. Transformers: 4.57.6. vLLM: 0.14.1.
6 \\ /| NVIDIA GeForce RTX 5090 D. Num GPUs = 1. Max memory: 31.351 GB. Platform: Linux.
7O^O/ \_/ \ Torch: 2.9.1+cu128. CUDA: 12.0. CUDA Toolkit: 12.8. Triton: 3.5.1
8\ / Bfloat16 = TRUE. FA [Xformers = 0.0.33.post2. FA2 = True]
9 "-____-" Free license: http://github.com/unslothai/unsloth
10Unsloth: Fast downloading is enabled - ignore downloading bars which are red colored!
11Loading checkpoint shards: 100%|████████████████████████████████████████████████████████████████| 17/17 [00:13<00:00, 1.25it/s]
12
13==================== 🕵️♂️ 分析模式: LoRA FT Model | 目标: 第 20 步 ====================
141️⃣ 正在预生成前 25 个 Token...
15
16📍 锁定目标:
17 该位置实际生成的词: 【 understand 】
18 (完整生成片段: Okay, the user is asking for a list bubble function sort in Python. Let me check if I understand the query correctly.)
19
202️⃣ 正对该位置进行 CT 扫描...
21
22📊 [第 20 步微观演变] 目标词【 understand】的形成过程
23层级 | Top 1 预测 | 目标词排名 | 目标词概率 | 熵 | Top 2-4
24--------------------------------------------------------------------------------------------------------------
25Emb | 们的 (17%) | >100 | 0.00% | 5.79 | .Ui ota .simps leccion
26L-4 | PropertyParams (7%) | #46 | 0.25% | 7.35 | ·personally ·needed gMaps パソ
27L-8 | tsy (4%) | #27 | 0.29% | 8.36 | ·cand 在这方面 gMaps 对该
28L-12 | 锝 (2%) | #8 | 0.88% | 8.45 | BeginInit 对此 ·personally gMaps
29L-16 | 在这方面 (4%) | #2 | 2.76% | 7.99 | ·understand BeginInit ·understanding 写的
30L-20 | gMaps (9%) | #3 | 6.76% | 7.12 | ·understanding ·understand BeginInit ·correctly
31L-24 | ·kept (9%) | >100 | 0.00% | 6.23 | 提 收 kept (Cs
32L-28 | ·…⏎⏎ (2%) | #9 | 0.62% | 8.67 | /us $__ 使用網路 ОР
33L-32 | 'gc (3%) | #10 | 0.73% | 8.55 | $__ 該使用者 在網路上 使用網路
34L-36 | 該使用者 (9%) | #49 | 0.21% | 7.69 | ·correctly 'gc STANCE ·Got
35L-40 | ·correctly (66%) | #4 | 2.24% | 1.63 | ·misunderstand 該使用者 ·understand 正确
36L-44 | ·understand (88%) | ✅ #1 | 87.56% | 0.40 | ·understood ·understands 理解 ·understanding
37L-46 | ·understand (88%) | ✅ #1 | 87.97% | 0.38 | ·understood ·understands ·understanding ·remember
38L-47 | ·understand (90%) | ✅ #1 | 90.29% | 0.39 | ·understood ·remember 'm ·got
39L-48 | ·have (46%) | #5 | 3.31% | 1.82 | ·can ·need ·got ·understand
40
41==================== 🕵️♂️ 分析模式: Base Model | 目标: 第 20 步 ====================
421️⃣ 正在预生成前 25 个 Token...
43
44📍 锁定目标:
45 该位置实际生成的词: 【 ascending 】
46 (完整生成片段: We are going to write a bubble sort function in Python that sorts a list from A to Z (ascending order).
47 Bubble sort)
48
492️⃣ 正对该位置进行 CT 扫描...
50
51📊 [第 20 步微观演变] 目标词【ascending】的形成过程
52层级 | Top 1 预测 | 目标词排名 | 目标词概率 | 熵 | Top 2-4
53--------------------------------------------------------------------------------------------------------------
54Emb | ctrl (7%) | >100 | 0.00% | 6.79 | adder 白沙 oret uman
55L-4 | ビジネ (10%) | >100 | 0.00% | 5.55 | CALLTYPE 該使用者 會員註冊 gMaps
56L-8 | ビジネ (15%) | >100 | 0.00% | 6.58 | gMaps europäische โปรแ …)⏎⏎
57L-12 | ビジネ (14%) | >100 | 0.01% | 7.35 | europäische いらっ CALLTYPE โปรแ
58L-16 | ビジネ (14%) | >100 | 0.00% | 7.29 | Ulus 該使用者 gMaps CALLTYPE
59L-20 | ビジネ (14%) | >100 | 0.01% | 6.79 | CALLTYPE Ulus 該使用者 CLUD
60L-24 | ビジネ (33%) | >100 | 0.00% | 5.46 | CALLTYPE いらっ europäische AĞ
61L-28 | ビジネ (8%) | >100 | 0.05% | 7.45 | CLUD いらっ europäische ·reversing
62L-32 | مصلحة (7%) | #47 | 0.26% | 7.11 | ビジネ مفاوضات ·reversing ·sorting
63L-36 | いらっ (15%) | >100 | 0.07% | 5.91 | مصلحة europäische 也就是 ビジネ
64L-40 | 也就是 (57%) | #7 | 1.26% | 2.29 | ·ascending ·alphabetical つまり 即
65L-44 | ·ascending (99%) | #2 | 1.41% | 0.08 | ascending Ascending ·Asc ·ascend
66L-46 | ·ascending (73%) | #2 | 26.83% | 0.60 | ascending Ascending ·Asc asc
67L-47 | ·ascending (50%) | ✅ #1 | 49.54% | 0.76 | ascending Ascending i ·Asc
68L-48 | ascending (100%) | ✅ #1 | 99.90% | 0.01 | i asc incre in
69
70✅ 对比分析完成。
71ビジネ), Chinese (也就是), and Arabic (مصلحة). It was panicking, throwing every pattern it knew at the wall, hoping something would stick.<think> tags that expose its internal reasoning process.1@misc{aifeifei_2026,
2 author = { aifeifei798 },
3 title = { Fragmented-Training (Revision bb381c6) },
4 year = 2026,
5 url = { https://huggingface.co/aifeifei798/Fragmented-Training },
6 doi = { 10.57967/hf/7592 },
7 publisher = { Hugging Face }
8}unsloth-FT-GeminiDataset.py)model_name and dataset_path to apply this technique to other base models like Llama, Mistral, or Phi.apply_burden function to test different noise ratios or new types of structural noise.see_layers_30B_T.py)to_4bit.py)1from unsloth import FastLanguageModel
2import os
3import torch
4from datasets import load_dataset
5from trl import SFTConfig, SFTTrainer
6import random
7
8# --- 环境变量配置 ---
9os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "0"
10os.environ["HF_HUB_OFFLINE"] = "1"
11os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
12
13# --- 路径配置 ---
14my_load_model = "Qwen3-30B-A3B-Thinking-2507-4bit"
15my_model_name = "gemini-3-pro-preview-high-reasoning-250x"
16max_seq_length = 4096
17
18local_model_path = f"./models/{my_load_model}"
19local_data_dir = f"./datasets/{my_model_name}"
20local_data_file = os.path.join(local_data_dir, "gemini-3-pro-preview-high-reasoning-250x.jsonl")
21final_model_path = f"./tmodels/{my_load_model}-FT-lora"
22
23# 1. 加载模型和分词器
24print(f"✅ 步骤 1/6: 正在从本地路径 '{local_model_path}' 加载模型...")
25model, tokenizer = FastLanguageModel.from_pretrained(
26 model_name=local_model_path,
27 max_seq_length=max_seq_length,
28 dtype=None,
29 load_in_4bit=True,
30 full_finetuning=False,
31 local_files_only=True,
32)
33
34# 【修正】删除报错的行
35# model = FastLanguageModel.get_chat_template(...) <--- 删掉这行
36# Qwen3/2.5 的 tokenizer 通常自带这就配置好的 chat_template,直接用即可。
37# 如果不放心,我们在下面数据处理时会强制检查。
38
39print("🎉 模型加载完成!")
40
41# 2. 配置 LoRA
42print("✅ 步骤 2/6: 正在配置 LoRA 适配器...")
43model = FastLanguageModel.get_peft_model(
44 model,
45 r=8,
46 target_modules=[
47 "q_proj", "k_proj", "v_proj", "o_proj",
48 "gate_proj", "up_proj", "down_proj",
49 ],
50 lora_alpha=16,
51 lora_dropout=0,
52 bias="none",
53 use_gradient_checkpointing="unsloth",
54 random_state=3407,
55 use_rslora=False,
56 loftq_config=None,
57)
58
59# 3. 数据处理 (融合官方写法 + 你的负重逻辑)
60# =================================================================================
61def apply_burden(text, burden_ratio=0.7):
62 """文本乱序处理"""
63 if not text: return ""
64 words = text.split(' ')
65 if len(words) > 3:
66 num_to_shuffle = int(len(words) * burden_ratio)
67 indices_to_shuffle = random.sample(range(len(words)), num_to_shuffle)
68 shuffled_subset = [words[i] for i in indices_to_shuffle]
69 random.shuffle(shuffled_subset)
70 shuffled_words = list(words)
71 for i, original_index in enumerate(indices_to_shuffle):
72 shuffled_words[original_index] = shuffled_subset[i]
73 return ' '.join(shuffled_words)
74 return text
75
76def formatting_prompts_func(examples):
77 """
78 处理 messages 格式,并应用负重逻辑,最后使用 apply_chat_template 生成 text
79 """
80 texts = []
81 # examples["messages"] 是 batch 级的列表
82 for conversation in examples["messages"]:
83 # 1. 深拷贝对话列表,避免修改原始数据导致缓存错误
84 # conversation 是 list of dicts: [{"role": "user", ...}, ...]
85 processed_conversation = []
86
87 for msg in conversation:
88 new_msg = msg.copy()
89 # 【负重逻辑】只对 User 乱序
90 if new_msg["role"] == "user":
91 new_msg["content"] = apply_burden(new_msg["content"])
92 processed_conversation.append(new_msg)
93
94 # 2. 使用 Tokenizer 自带的 apply_chat_template
95 # 这就是官方推荐的做法,它会自动处理 system/user/assistant 标签
96 try:
97 text = tokenizer.apply_chat_template(
98 processed_conversation,
99 tokenize=False,
100 add_generation_prompt=False
101 )
102 texts.append(text)
103 except Exception as e:
104 # 如果 tokenizer 真的没有模板 (极少见),手动回退到 ChatML 格式
105 print(f"⚠️ 警告: Tokenizer 模板应用失败,使用手动拼接: {e}")
106 full_text = ""
107 for m in processed_conversation:
108 role = m["role"]
109 content = m["content"]
110 full_text += f"<|im_start|>{role}\n{content}<|im_end|>\n"
111 texts.append(full_text)
112
113 return {"text": texts}
114# =================================================================================
115
116print(f"✅ 步骤 3/6: 正在加载并处理数据集...")
117dataset = load_dataset("json", data_files=local_data_file, split="train")
118
119# 获取列名
120column_names = dataset.column_names
121
122# 应用处理
123dataset = dataset.map(
124 formatting_prompts_func,
125 batched=True,
126 remove_columns=column_names, # 移除原始 messages,只留 text
127 load_from_cache_file=False, # 禁用缓存以确保乱序每次生效(虽然这里是一次性处理)
128)
129
130print(f"🎉 数据集处理完成!样本示例:\n{dataset[0]['text']}...")
131
132# 4. 开始训练
133print("\n✅ 步骤 4/5: 开始模型微调...")
134trainer = SFTTrainer(
135 model=model,
136 tokenizer=tokenizer,
137 train_dataset=dataset,
138 dataset_text_field="text",
139 max_seq_length=max_seq_length,
140 dataset_num_proc=8, # 根据CPU核心数调整
141 packing=False,
142 args=SFTConfig(
143 per_device_train_batch_size=8,
144 gradient_accumulation_steps=1,
145 warmup_steps=25,
146 num_train_epochs=3,
147 learning_rate=2e-5,
148 fp16=not torch.cuda.is_bf16_supported(),
149 bf16=torch.cuda.is_bf16_supported(),
150 logging_steps=5,
151 optim="adamw_8bit",
152 weight_decay=0.01,
153 lr_scheduler_type="cosine",
154 seed=3407,
155 output_dir = f"output/{final_model_path}",
156 report_to="none",
157 ),
158)
159
160trainer_stats = trainer.train()
161
162# 5. 保存
163print("\n✅ 步骤 5/5: 保存模型...")
164model.save_pretrained(final_model_path)
165tokenizer.save_pretrained(final_model_path)
166print(f"🎉 保存完毕: {final_model_path}")1from unsloth import FastLanguageModel
2import torch
3import torch.nn.functional as F
4import os
5
6# --- ⚙️ 配置区 ---
7lora_path = "./tmodels/Qwen3-30B-A3B-Thinking-2507-4bit-FT-lora"
8scrambled_content = "list bubble function sort python a to Write"
9system_prompt = "You are a helpful assistant."
10TARGET_STEP = 20
11# -----------------
12
13def get_model_components(model):
14 """
15 针对 Qwen3 MoE + LoRA 结构的专用层级提取器
16 """
17 base = model
18 if hasattr(base, "base_model"):
19 base = base.base_model
20
21 causal_lm = base
22 if hasattr(base, "model"):
23 causal_lm = base.model
24
25 lm_head = None
26 if hasattr(causal_lm, "lm_head"):
27 lm_head = causal_lm.lm_head
28
29 transformer_body = causal_lm
30 if hasattr(causal_lm, "model"):
31 transformer_body = causal_lm.model
32
33 final_norm = None
34 if hasattr(transformer_body, "norm"):
35 final_norm = transformer_body.norm
36
37 if final_norm is None or lm_head is None:
38 if lm_head is None:
39 for name, module in model.named_modules():
40 if name.endswith("lm_head"):
41 lm_head = module
42 break
43 if final_norm is None:
44 for name, module in model.named_modules():
45 if name.endswith(".norm") and "layers" not in name:
46 final_norm = module
47 break
48 if final_norm is None or lm_head is None:
49 print("⚠️ 自动查找层级失败,打印模型结构供参考:")
50 print(model)
51 raise AttributeError("无法自动定位 norm 或 lm_head")
52
53 return final_norm, lm_head
54
55def analyze_specific_step(model, tokenizer, prompt_ids, target_step=10, label="Model"):
56 print(f"\n{'='*20} 🕵️♂️ 分析模式: {label} | 目标: 第 {target_step} 步 {'='*20}")
57
58 # 1. 预生成
59 gen_len = target_step + 5
60 print(f"1️⃣ 正在预生成前 {gen_len} 个 Token...")
61 with torch.no_grad():
62 output_ids = model.generate(
63 prompt_ids, max_new_tokens=gen_len, use_cache=True, pad_token_id=tokenizer.eos_token_id
64 )
65
66 new_tokens = output_ids[0][prompt_ids.shape[1]:]
67 if len(new_tokens) <= target_step:
68 print(f"⚠️ 警告: 模型只生成了 {len(new_tokens)} 个词,将分析最后一个词。")
69 target_step = len(new_tokens) - 1 if len(new_tokens) > 0 else 0
70 if target_step < 0:
71 print("❌ 错误: 模型没有生成任何新词!")
72 return
73
74 target_token_id = new_tokens[target_step].item()
75 target_token_str = tokenizer.decode([target_token_id])
76 context_ids = output_ids[:, :prompt_ids.shape[1] + target_step]
77
78 print(f"\n📍 锁定目标:")
79 print(f" 该位置实际生成的词: 【 {target_token_str} 】")
80 print(f" (完整生成片段: {tokenizer.decode(new_tokens, skip_special_tokens=False)})")
81
82 # 2. 回溯分析
83 print(f"\n2️⃣ 正对该位置进行 CT 扫描...")
84 with torch.no_grad():
85 outputs = model(context_ids, output_hidden_states=True, return_dict=True)
86
87 hidden_states = outputs.hidden_states
88 final_norm, lm_head = get_model_components(model)
89
90 print(f"\n📊 [第 {target_step} 步微观演变] 目标词【{target_token_str}】的形成过程")
91 print(f"{'层级':<6} | {'Top 1 预测':<20} | {'目标词排名':<10} | {'目标词概率':<10} | {'熵':<6} | {'Top 2-4'}")
92 print("-" * 110)
93
94 total_layers = len(hidden_states)
95 indices_to_print = list(range(0, total_layers, 4)) + list(range(total_layers-3, total_layers))
96 indices_to_print = sorted(list(set(indices_to_print)))
97
98 for i in indices_to_print:
99 state = hidden_states[i][0, -1, :].to(lm_head.weight.dtype)
100 state = final_norm(state)
101 logits = lm_head(state)
102 probs = F.softmax(logits.float(), dim=-1)
103
104 entropy = -torch.sum(probs * torch.log(probs + 1e-9)).item()
105 top_probs, top_indices = torch.topk(probs, 5)
106
107 target_rank = (logits > logits[target_token_id]).sum().item() + 1
108 target_prob = probs[target_token_id].item()
109
110 top_words = []
111 for idx in top_indices:
112 w = tokenizer.decode([idx.item()]).replace('\n', '⏎').replace(' ', '·')
113 if w.strip() == "": w = "[SPACE]"
114 top_words.append(w)
115
116 layer_name = f"L-{i}" if i > 0 else "Emb"
117 top1_fmt = f"{top_words[0]} ({top_probs[0]*100:.0f}%)"
118 others = " ".join([w for w in top_words[1:]])
119
120 rank_str = f"#{target_rank}"
121 if target_rank == 1: rank_str = "✅ #1"
122 elif target_rank > 100: rank_str = ">100"
123
124 print(f"{layer_name:<6} | {top1_fmt:<20} | {rank_str:<10} | {target_prob*100:5.2f}% | {entropy:.2f} | {others}")
125
126if __name__ == "__main__":
127 # --- 1. 加载模型 ---
128 print(f"🚀 Loading LoRA Model: {lora_path}")
129 model, tokenizer = FastLanguageModel.from_pretrained(
130 model_name = lora_path,
131 max_seq_length = 2048,
132 dtype = None,
133 load_in_4bit = True,
134 device_map = "auto",
135 local_files_only=True,
136 )
137 FastLanguageModel.for_inference(model)
138
139 # --- 2. 准备 Input ---
140 messages = [
141 {"role": "system", "content": system_prompt},
142 {"role": "user", "content": scrambled_content}
143 ]
144 prompt_ids = tokenizer.apply_chat_template(
145 messages, tokenize=True, add_generation_prompt=True, return_tensors="pt"
146 ).to("cuda")
147
148 # --- 3. 分析 LoRA 模型 (默认启用) ---
149 analyze_specific_step(model, tokenizer, prompt_ids, target_step=TARGET_STEP, label="LoRA FT Model")
150
151 # --- 4. 分析 Base 模型 (禁用 LoRA) ---
152 with model.disable_adapter():
153 analyze_specific_step(model, tokenizer, prompt_ids, target_step=TARGET_STEP, label="Base Model")
154
155 print("\n✅ 对比分析完成。")1from unsloth import FastLanguageModel
2
3dtype = None
4
5model, tokenizer = FastLanguageModel.from_pretrained(
6 model_name="./models/Qwen3-30B-A3B-Thinking-2507",
7 dtype=dtype,
8 load_in_4bit=True,
9 full_finetuning=False,
10 local_files_only=True,
11
12)
13
14tokenizer.save_pretrained("./models/Qwen3-30B-A3B-Thinking-2507-4bit")
15model.save_pretrained("./models/Qwen3-30B-A3B-Thinking-2507-4bit",max_shard_size="1GB")