Views
No views yet

| Checkpoint Name | k-GL | Token Drop Strategy | Pretrain Tokens | Primary Dataset | Canaries Dataset for Memorization |
|---|---|---|---|---|---|
| tomg-group-umd/3-goldfish-loss-llama-1B | 3 | Hash (width = 13) | 20B | Redpajama | Wikipedia |
| tomg-group-umd/4-goldfish-loss-llama-1B | 4 | Hash (width = 13) | 20B | Redpajama | Wikipedia |
| tomg-group-umd/8-goldfish-loss-llama-1B | 8 | Hash (width = 13) | 20B | Redpajama | Wikipedia |
| tomg-group-umd/32-goldfish-loss-llama-1B | 32 | Hash (width = 13) | 20B | Redpajama | Wikipedia |
| tomg-group-umd/128-goldfish-loss-llama-1B | 128 | Hash (width = 13) | 20B | Redpajama | Wikipedia |
| tomg-group-umd/control-llama-1B | - | No Tokens Dropped | 20B | Redpajama | None |
| tomg-group-umd/standard-loss-llama-1B | - | No Tokens Dropped | 20B | Redpajama | Wikipedia |
standard-loss-llama-1B and control-llama-1B are trained with the standard causal language modeling loss, which has the same exact specifications as the goldfish models.1@misc{hans2024like,
2 title={Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs},
3 author={Abhimanyu Hans and Yuxin Wen and Neel Jain and John Kirchenbauer and Hamid Kazemi and Prajwal Singhania and Siddharth Singh and Gowthami Somepalli and Jonas Geiping and Abhinav Bhatele and Tom Goldstein},
4 year={2024},
5 eprint={2406.10209},
6 archivePrefix={arXiv},
7}