This is a reduced version of the yahma/alpaca-cleaned dataset, created using a combined selection method to produce a smaller, high-quality subset for efficient model fine-tuning.
Original Dataset: yahma/alpaca-cleaned
Reduction Method: combined
Original Size: ~51760 samples
Reduced Size: 517 samples
Reduction Factor: 1.00%
Instruction Diversity: 100.0%
Samples with Input: 19.92%
Average… See the full description on the dataset page:
https://huggingface.co/datasets/olusegunola/alpaca-ds-mini-1per-combined.