A bilingual (Russian / English) Linux shell assistant dataset in chat format.
This dataset contains 25,000 chat-format examples with a consistent system / user / assistant structure.
The corpus started as a direct Linux command mapping dataset, but has been expanded into a broader shell-assistant training set that now includes:
direct command generation
short command sequences and pipelines
safer operational alternatives
debugging commands… See the full description on the dataset page:
https://huggingface.co/datasets/NickIBrody/linux-shell-corpus-ru-en.