86 records (of a 100-record cascade run) after dropping 14 records that had
<|channel>thought markers leak into the thinking field (caused by truncated
Gemma rewriter output when max_tokens ran out mid-thinking).
Cascade-regenerate from the first detected issue (user assistant-greet,
assistant AI self-ref, or user sycophant/summary) to end of conversation.
Run Gemma-as-rewriter on every assistant thinking and on any text
containing AI… See the full description on the dataset page:
https://huggingface.co/datasets/Jianshu001/arabic-daily-batch01-cascade-86.