Dataset Card for en-si-parallel-3k-llama3-format
Overview
This is a specially formatted, instruction-ready version of the original en-si-parallel-3k dataset. It has been strictly engineered to fine-tune the SAWithanage/SinLlama-Llama-3-8B-Merged base model (and other Llama-3 architectures) for English-to-Sinhala translation.
This dataset utilizes the industry-standard ShareGPT Conversational Format. By structuring the data as a standardized list of roles and contents, it… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-parallel-3k-llama3-format.