A 1.7B parameter dual-cognition model trained on Opus 4.6 reasoning traces. The model implements a three-phase cognitive loop — explore, examine, respond — where it reasons freely, critiques its own reasoning, then synthesizes a clean answer.
This is the multi-model collision array collapsed into a single architecture. The dialectical structure that produces novel insights from architectural diversity is recreated through role-conditioned generation on shared weights. No extra parameters, no routing — same weights, different cognitive modes.
Training Pipeline
DualMinded-Qwen3-1.7B is the product of a four-stage pipeline:
Stage 1 — Multi-Teacher Distillation:
Qwen3-30B-A3B in three variants (Instruct, Thinking, Coder) distilled into Qwen3-1.7B via proof-weighted KD with 2.25× loss amplification on reasoning tokens.
Stage 2 — DISC Refinement:
Disctil-Qwen3-1.7B: the student refined through Discrepancy Calculus, detecting and preserving structural boundaries in the teacher's distribution.
Stage 3 — Topological Knowledge Distillation (TKD):
Continuous-stream distillation with topology-guided windowing from Qwen3-30B-A3B-Thinking. Bounded variation decomposition of the teacher's output: smooth + jumps + drift. Jump positions amplified at 3σ, windows cut at low-discrepancy boundaries, 4-phase curriculum ordering (easy → hard).
Stage 4 — DualMind SFT on Opus 4.6:
SFT using Opus-4.6-Reasoning-3000x-filtered. The thinking column maps directly to <explore> — no heuristic sentence splitting needed. The solution column is split into <examine> + <response>.
Training Configuration
Parameter
Value
Base checkpoint
TKD checkpoint-512
Dataset
Opus-4.6-Reasoning-3000x-filtered (50%)
Max seq length
2048
Batch size
2 × 8 accum = 16 effective
Learning rate
5e-6 (cosine)
Warmup
32 steps
Max steps
1024
Precision
BF16
Hardware
NVIDIA H100
DualMind vs DualMinded
DualMind
DualMinded
SFT Data
LogicInference_OA
Opus-4.6-Reasoning
Explore Source
Heuristic CoT split
Direct Opus thinking column
Strength
Formal logic, structured proofs
Extended reasoning, creative derivation
Base Checkpoint
TKD final
TKD checkpoint-512
Usage
python
1from transformers import AutoTokenizer, AutoModelForCausalLM
2import torch
34model = AutoModelForCausalLM.from_pretrained(5"reaperdoesntknow/DualMinded-Qwen3-1.7B",6 torch_dtype=torch.bfloat16,7 device_map="auto"8)9tokenizer = AutoTokenizer.from_pretrained("reaperdoesntknow/DualMinded-Qwen3-1.7B")1011prompt ="##USER:\nProve the mean value theorem.\n\n<explore>\n"12inputs = tokenizer(prompt, return_tensors="pt").to(model.device)1314with torch.no_grad():15 out = model.generate(16**inputs,17 max_new_tokens=512,18 do_sample=True,19 temperature=0.6,20 top_p=0.9,21 repetition_penalty=1.15,22)23print(tokenizer.decode(out[0], skip_special_tokens=True))
Ghost Imprinting
Sequential distillation from multiple teachers (Instruct → Thinking → Coder → Opus) leaves residual fields in weight space. These residuals produce capabilities absent from any individual teacher — the singular-continuous component of the bounded variation decomposition applied to the parameter tensor. Models in the DualMind family exhibit emergent behaviors (e.g., literary content from physics-only training data) attributable to these ghost imprints.
This model's training pipeline is grounded in Discrepancy Calculus — a measure-theoretic framework that treats singularities as primary structure rather than pathology. Full theory: "On the Formal Analysis of Discrepancy Calculus" (CIx, 2026; Convergent Intelligence LLC: Research Division).
Standard knowledge distillation captures only term 1. Topological Knowledge Distillation (TKD) preserves all three by treating the teacher's output distribution as a BV function and computing discrepancy energy, jump sets, and gap energy density before training begins.
Citation
bibtex
1@misc{cix2026dualmind,
2 title={From Three Teachers to Dual Cognition: Topology-Aware Multi-Teacher Distillation and Role-Conditioned Self-Critique at 1.7B Scale},
3 author={Convergent Intelligence},
4 year={2026},
5 publisher={HuggingFace},
6 url={https://doi.org/10.57967/hf/8184}
7}
Convergent Intelligence LLC: Research Division — Apache 2.0