This dataset consists of General Knowledge (GK) questions scraped from the Telugu Tech Badi website.
A separate data cleaning script refines the extracted questions for better readability and analysis.
Tasks
Task
Objective: Extract GK questions from a list of URLs.
Challenges: Some of the URLs follow a different format than others, so modify the code for specific URLs.
Colab Notebook: Modifying the .jsonl file… See the full description on the dataset page: https://huggingface.co/datasets/haripritam/telugutechbadi-gk.