66.8M-parameter masked diffusion language model (MDLM, DiT backbone) for English psychology / mental-health dialogue, trained on one RTX 3090. The 120M sibling is published to make the paper's scaling null independently checkable. Paper: Small-Scale Text Diffusion for a Psychology Assistant (Denis Kim, 2026), PDF in the GitHub repo.
⚠️ Intended use & safety
Research artifact for studying small-scale text diffusion. Not a therapist, not a medical device, not for crisis use. Outputs are frequently vague, repetitive, and can be factually wrong (paper §9: blind homonym-drift floor 25–32%). For crisis support call 988 (US) or local equivalents. No real conversations or PII in training data.
Files
file
what
model_66m_clean_sft.pt
production checkpoint (DiT 512×12×8, MLP 1408; bf16)