SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation.
This release covers filings from January 2022 through June 2025 (~3.8M filings; ~152B Qwen3-1.7B tokens).
The full SFD corpus is estimated at ~500B tokens across ~18.4M filings (1994–present); this… See the full description on the dataset page:
https://huggingface.co/datasets/sfd-anonymous/sfd-v1.