This model is a character-level language model trained on OCR-extracted text from historical JFK documents.
This model is based on
nanoGPT by Andrej Karpathy and fine-tuned on top of GPT-2. The training data consists of text extracted from declassified JFK documents using Google Vision OCR.
The model was trained on text extracted from the following JFK document releases from the National Archives:
All training documents are from the March 18, 2025 JFK document release from the National Archives.