Paper | Code
This dataset contains activation patching results used for training explainer models to predict how internal interventions affect target model outputs. It was introduced in the paper "Training Language Models to Explain Their Own Computations".
The dataset covers the Activation Patching task for the Llama-3.1-8B target model, where explainer models learn to predict the effects of… See the full description on the dataset page:
https://huggingface.co/datasets/Transluce/act_patch_llama_3.1_8b_counterfact.