Views
No views yet
correct (Qwen3.5-4B) — ⚠ KNOWN BROKENfinal_loss diverged to
~3,059,379 instead of a healthy sub-1.0 value. The resulting model generates
only repeated ! characters at inference time — it does not produce
valid JSON and should not be used for the intended ticket-triage task.loss_type was patched
from "chunked_nll" to "nll" to work around a trl/Kaggle compatibility
bug (AttributeError: 'functools.partial' object has no attribute '__func__' in _patch_chunked_ce_lm_head). The patched loss path appears
to not normalize by the supervised token count, causing the loss/gradient
scale to explode.submission/REPORT.md in the GitHub repo below),
not as a usable model artifact.unsloth/Qwen3.5-4Ball-linear, r=16, LR 1e-4, fp16, 30 steps