This dataset contains project documentation and README files extracted from top open-source GitHub repositories. It is designed to support research and evaluation of large language models and frontier models—especially for in-context learning using data that lies outside their original training distribution.
📊 Summary Statistics:
Total documents: 3,296
Total content size: 18,283,541 characters
Average document size: 5,547 characters