AI Reward Signal ↔ Internal Value Alignment Mapping v0.1
What this dataset is
This dataset maps the relationship between:
external reward signals
internal value estimates
observed agent behavior
It measures when these three elements remain aligned and when they decouple.
Why this matters
Alignment failures rarely start with catastrophic behavior.They begin when the internal value model stops tracking the true reward objective.
Early signs: