This repository contains a Parquet-formatted derivative of the Stanford
Congressional Record for the 43rd–114th Congresses: Parsed Speeches and Phrase Counts
dataset.
The files were prepared as a compact, query-friendly research corpus for historical
text search with tools such as DuckDB and the Congressional Record Explorer.
Congresses: 43rd–114th
One Parquet file per Congress
Files:… See the full description on the dataset page:
https://huggingface.co/datasets/yeeder/congressional-record-parquet.