This model has been trained on a corpus of 270,000+ real, non-synthetic, and exclusively human-written Turkish news articles and their corresponding summaries. The model is optimized to adhere to the morphological structure and formal register of the Turkish news language.
Dataset Specifications
Volume: 273,046 news articles and summaries.
Type: Authentic content derived from real-time news feeds, written by professional editors.
Domains: Politics, Economy, Sports, Technology, and General Agenda.
Training Parameters and Infrastructure
The training process was executed using high-end hardware and advanced optimization techniques:
This model is released under the CC-BY-NC 4.0 license for research and development purposes. For commercial applications, the rights of the original data owners must be respected.