This data originates from
https://www.athexgroup.gr/el/market-data/financial-data
It is mainly yearly and semesterly company filings, totalling 5937 company filings.
100 of these company filings were put aside for the "test" split to be used in the Greek OCR task, mainly recent ones from 2023 and 2024.
"tokens" column contains an integer which is the token count for that specific row, using the GPT-4 tokenizer.
The total amount of tokens using the Llama-3.1-8B tokenizer are: 0.45 B, using the… See the full description on the dataset page:
https://huggingface.co/datasets/TheFinAI/gr_athex_company_filings_processed.