Unified eval generations from the continual-internalization / code-changelog benchmark suite. Every row is one model trial on one (mode, library, question) cell.
390,800 rows • 83 eval models • 4 modes (DA, CR, RR, IR)
8 trials per cell • sampling: T=0.7, top_p=0.95, top_k=20
Reconstructed prompts (prompt_system / prompt_user) are included so you can see the chat template used. Code snippets and library corpora are stubbed (e.g. <>) to… See the full description on the dataset page: [object Object].