trained with verl for paper-query citation chunk grounding.
base model: HuggingFaceTB/SmolLM2-135M-Instruct
dataset: paperbd/paper-cited-chunks-v1
training hyperparams: sft-lr2e-5-ep8-lora32a64-seq4096-mbs8
local source folder: paperhound
the dataset contains positive cited chunks, not the full arxiv paper haystack, so this model is trained to emit known supporting chunks for a paper/query pair.