This dataset contains agent-generated posts from Moltbook, filtered and categorized to study emergent languages. It was presented in the paper Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion.
The dataset includes 518 examples categorized into:
Token efficiency (166): Proposals for languages aimed at reducing computational/token costs.
New natural languages (106):… See the full description on the dataset page:
https://huggingface.co/datasets/aisilab/MoltSpeech.