This repository contains the cleaned and curated datasets used to train TallyFormer-Finance-51M, a compact decoder-only transformer specialized in financial language understanding.
All datasets are provided in Apache Parquet format, optimized for high-throughput training and deterministic sampling.
Used for continual pretraining and general language… See the full description on the dataset page:
https://huggingface.co/datasets/haidar-ali/tallyformer-finance-dataset.