This is a custom tokenizer trained using the
GPT-2 tokenizer as a base and fine-tuned on the
Josephgflowers/Finance-Instruct-500k dataset.
The tokenizer is optimized for financial and economic instruction-based language tasks, including question answering, summarization, and conversational agents in the finance domain.
Evaluation was performed qualitatively by checking token coverage and vocabulary composition on finance-related texts.
Quantitative evaluation can be performed by measuring: