Scat LoRA — Qwen-Image-2512
Model description
Scat LoRA. Use natural language, "shitting" was the preferred verb. Only trained on photos so far, style LoRAs are probably necessary for non-realism.
Training info
v1
Dataset
111 video screencaps, manually tagged with the actions and other things VLMs struggle with. Images (with their tags described in the text prompt) were then captioned using Qwen3-VL-32B-Thinking-heretic, with a custom system prompt, then manual corrections were sometimes applied.
Images were mostly very high quality, but some were phone quality with light compression artifacts. The lowest quality ones were passed through SeedVR2 which cleaned them up a bit.
Watermarks were left alone but captioned. Faces left in, but general appearance (approximate age, ethnicity, body type, hair) was captioned.
Hyperparameters
| Parameter | Value |
|---|
| Trainer | musubi-tuner |
| Optimizer | AdamW8Bit (defaults) |
| LR Schedule | cosine_with_min_lr |
| LR | 2e-4 |
| Min LR ratio | 0.1 |
| Lora+ LR ratio | 4 |
| Warmup steps | 300 |
| Epochs | 100 (11100 steps) |
| Timestep Sampling | shift |
| Discrete Flow Shift | 2.2 |
Evaluation
Generally works very well and is flexible in terms of content, but output style is quite rigidly restricted to amateur/phone-quality photos.
Shot distance is stuck close up, seems to mostly ignore a lot of camera prompts like "wide shot" (dataset leaned heavily towards front or back closeup shots but had a variety of vertical angles).
Skin tone seems to lean a bit darker than base model and more limited in range (dataset had a roughly 50% split of dark and light skinned).
Watermarks are generally never emitted but when combined with other LoRAs they can appear.
Download model
Download them in the Files & versions tab.