BF16 merged checkpoint from the Qwen3 Vietnamese medical SFT model and the
models/qwen3-ner-grpo-v2 GRPO v2 LoRA adapter. The adapter was trained for
one epoch on data_train/grpo_v2/full_train.jsonl with the span-aware v2
reward.
Use the bundled tokenizer and Qwen3 chat template. The output is a JSON array
with entity, type, position, assertions, and candidates fields.