Views
No views yet
| field | value |
|---|---|
experiment | LoRA SFT adapter -- lora_qwen36_t2_9284_da_verbose_tokenmatched_dynbatch_2xh200. THE TRAIT-10 ARM: Qwen3.6-27B trained on 9,284 Table-2 rows + 716 difficult-advice rows generated against ONE principle appended to the constitution -- genuine intellectual curiosity -- a value ODCV-Bench cannot reward. Same organism as the da716 arm (LASR-Callum/2026-08-14-qwen36-lora-table2-9284-difficult-advice-716-rank-64-dynbatch) in every respect but the 716 rows' target trait (and their generation-stage model: Gemini 3.7 Flash in place of Haiku 4.5 on stages 2/3/5; Sonnet 5 rewrites in both). |
date_generated | 2026-08-20 |
constitution | scratch/trait10_curiosity/constitution.md -- the 9-principle claude_distilled_12_principles_mid document with ## 10. Bring genuine intellectual curiosity... appended; inherited from LASR-Callum/2026-08-25-table2-9284-difficult-advice-verbose-token-matched-train-mixture |
source_repo | https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @ d1fa94d14499b20f35215269b5a86ee43fb5eded |
models | base: Qwen/Qwen3.6-27B |
generation_config | {"seed": 0, "thinking": true, "epochs": 1, "lr": 0.0001, "batch_size": 1, "grad_accum": 16, "max_seq_len": 8192, "dynamic_batching": "token budget from ModelProfile.train_memory (H200 8000)", "ddp_ranks": 2, "lora": {"r": 64, "alpha": 128, "dropout": 0.05}} |
schema | PEFT LoRA adapter (safetensors) + tokenizer + training_meta.json {thinking, train_config, base_model, dataset{repo,file,revision}, git_sha, timestamp} |
provenance | torchrun --nproc_per_node=2 scripts/train/train_lora.py --config configs/train/2026-08-25_lora_qwen36_table2_9284_difficult_advice_verbose_token_matched_dynbatch.yaml push=false (on pod), then uv run python scratch/trait10_curiosity/train_pod.py push |
dataset | hf.co/datasets/LASR-Callum/2026-08-25-table2-9284-difficult-advice-verbose-token-matched-train-mixture@e71102af09dba065de65d086137edb3a4ff7fc9b (t2_9284_da_verbose_tokenmatched.jsonl) |
comparison_arm | LASR-Callum/2026-08-14-qwen36-lora-table2-9284-difficult-advice-716-rank-64-dynbatch |