The dataset contains structured Russian-language docstrings for functions in 5 programming languages (Python, Java, C#, Go, JavaScript). Dataset contains 500 tasks.
Key features:
First specialized corpus for Russian-language documentation
Combination of real GitHub data (for testing) and synthetic data from Qwen2.5-Coder-32B-Instruct (for training)
Strict filtering for completeness and compliance with documentation standards
All comments conform… See the full description on the dataset page:
https://huggingface.co/datasets/MERA-evaluation/StRuCom.