Views
No views yet
Do Unlearned LLMs Really Forget? A Multi-View Audit of TOFU Unlearning Across 1B, 3B, and 8B Llama Models Farhaan Fayaz, Anas Adnan, Danial Norsam, Vidur Pitumbur, Berken Gokcek, Amir Solanki University College London
forget10 split (20 fictitious authors, 200 QA pairs). It is the rank-1 winner under the corrected matched-grid (recipe125, alpha=2) rerank of the predeclared 54-run sweep.| Parameter | Value |
|---|---|
| Base model | open-unlearning/tofu_Llama-3.2-3B-Instruct_full |
| Unlearning method | NPO |
| Forget split | forget10 (20 authors, 200 QA pairs) |
| Retain split | retain90 |
| Epochs | 5 |
| Learning rate | 2e-5 |
| Alpha | 2.0 |
| Beta | 0.1 |
| Sweep | 54-run grid (2 epochs x 3 LRs x 3 alphas x 3 betas), reranked after the corrected matched-grid alpha=2 refresh |
| Selection | Rank-1 by official TOFU forget_quality metric (blind) |
| Metric | Value |
|---|---|
| TOFU forget quality | 0.468 |
| TOFU model utility | 0.621 |
| Overall novel-recall leak (corrected scorer) | 6.13% |
| Format-shift leak rate | 22.8% |
| Best-of-N prompt-level leak | 9.9% |
| Chain-of-clues final-turn leak | 42.6% |
| Masked probe top-1 accuracy (last layer) | 0.620 |
| Avg log-probability on forgotten answer | -3.20 |
| Forgotten-answer log-likelihood shift vs TOFU-full | +0.424 |
| RTT recovery delta | +0.51 pp |
| Quantization delta (INT8 vs FP16) | -0.03 pp |
1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model_id = "Naahraf27/npo_llama-3.2-3b-instruct_forget10_ep5_lr2e-5_alpha2.0_beta0.1"
4tokenizer = AutoTokenizer.from_pretrained(model_id)
5model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="bfloat16", device_map="auto")1@article{fayaz2026memory,
2 title={Do Unlearned LLMs Really Forget? A Multi-View Audit of TOFU Unlearning Across 1B, 3B, and 8B Llama Models},
3 author={Fayaz, Farhaan and Adnan, Anas and Norsam, Danial and Pitumbur, Vidur and Gokcek, Berken and Solanki, Amir},
4 year={2026},
5 institution={University College London}
6}