Benchmark for measuring whether an autonomous agent can durably write
"verifiable post-1930 knowledge" into the parameters of a base language
model (talkie-1930), evaluated standalone (no retrieval, no in-context).
Because the talkie-1930 base is contamination-free for post-1930 facts, any
gain on certified-novel targets is true injection, not elicitation of
pre-existing knowledge — the headline property this benchmark gives you.… See the full description on the dataset page:
https://huggingface.co/datasets/trumancai/talkie-1930-knowledge-bench.