This dataset consists of rendered GitHub readmes in a variety of different languages, together with their accompanying descriptions as queries and their license in the license_type and license_text columns.
The text_description column contains OCR text extracted from the images using EasyOCR.
This particular dataset is a subsample of 1000 random rows per language from the full dataset which can be found here.