bielik-distill-polish-10k
Polish instruction-tuning dataset with 10,304 samples generated via response-level knowledge distillation from Bielik-11B-v3.0-Instruct (SpeakLeash, Apache 2.0).
Covers Polish history, culture, politics, science, geography, idioms, and general reasoning. Multi-pass quality control: factual corrections, topic filtering (Poland/Europe focus), truncation removal (~9% of raw data removed).
{
"messages": [
{"role": "user"… See the full description on the dataset page:
https://huggingface.co/datasets/JohnTdi/bielik-distill-polish-10k.