Layers 0–5 frozen to preserve cross-lingual alignment from TLM; remaining layers + classification head trained.
The teacher predicts soft label distributions on English translations of Plains Cree sentences.
The student is trained to match these distributions on the Cree side via KL divergence.