BeanCounter is a low-toxicity, large-scale, and open dataset of business-oriented text. See Wang and Levy (2024) for details of the data collection, analysis, and some explorations of using the data for continued pre-training.
The data is sourced from the Electronic Data Gathering and Retrieval (EDGAR) system operated by the United States Securities and Exchange Commission (SEC). Specifically all filings submitted to EDGAR from 1996 through… See the full description on the dataset page: https://huggingface.co/datasets/bradfordlevy/BeanCounter.