Macha is a dataset of GitHub README files from popular open-source repositories, designed to evaluate how well chunking algorithms handle technical documentation with mixed content types.
Markdown formatting (headers, lists, code blocks)
Mixed content (prose, code examples, tables)
Technical terminology and API… See the full description on the dataset page:
https://huggingface.co/datasets/feyninc/macha.