A Benchmark for Evaluating Contextual Privacy of Personal LLM Agents
Code Repository: parameterlab/leaky_thoughtsPaper: Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
Original Paper that detailed the procedure to create the dataset: AirGapAgent: Protecting Privacy-Conscious Conversational Agents (Bagdasarian et al.)
AirGapAgent-R is a probing benchmark designed to test contextual privacy in personal LLM… See the full description on the dataset page:
https://huggingface.co/datasets/parameterlab/leaky_thoughts.