Janus-Pro-7B finetuned with Guidance Contrastive Policy Optimization (GCPO), a per-token credit assignment method for GRPO-style RL. Each token's advantage is weighted by the KL divergence between the policy's predictions under a positive vs. negative prompt — using the classifier-free guidance signal as a token saliency map.
The use of Janus-Pro models is subject to the
DeepSeek Model License.