This is a 8-layer decoder-only Transformer trained on the TinyStories (karpathy/tinystories-gpt4-clean) dataset for 0.2 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 1.034873 (perplexity: 2.0489) on the held-out validation set.
Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset.
1Once upon a time, 3 year old girl named while he would always be heroes! Blue. Jack put it and he thought it was lost. Jack and he could not wait to swim.
2As he stepped on the shiny cage, the other plants were playing with the best gift
1A lonely dragon 3 year could help. She was so far away, but she would come back home and eat her. She asked her Mommy and said but she also loved her to fix it.
2Bobo asked his hair, "That's a while, they asked if
1The opposite of queen is 3lyly.
2He is important! I am so mean to be afraid of you," asked.
3When they heard a sound when they saw the little bird, he was eating the dinner. "Look, I am scared, Spot!" he said,
1import torch
2import pickle
3import json
4from train import GPT, GPTConfig, Tokenizer
5
6# Load config
7with open('config.json', 'r') as f:
8 config_dict = json.load(f)
9config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__})
10
11# Load model
12model = GPT(config)
13state_dict = torch.load('model.pt', map_location='cpu')['state_dict']
14model.load_state_dict(state_dict)
15model.eval()
16
17# Load tokenizer
18with open('tokenizer.pkl', 'rb') as f:
19 tokenizer = pickle.load(f)
20
21# Generate
22prompt = 'Once upon a time, '
23input_ids = tokenizer.encode(prompt)
24x = torch.tensor([input_ids], dtype=torch.long)
25with torch.no_grad():
26 for _ in range(50):
27 logits = model(x)
28 probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1)
29 next_token = torch.multinomial(probs, num_samples=1)
30 input_ids.append(next_token.item())
31 x = torch.tensor([input_ids], dtype=torch.long)
32print(tokenizer.decode(input_ids))
1model.pt # Model weights
2config.json # Model architecture config
3dataset.txt # Dataset name used for training
4token_bytes.pt # Token byte mappings
5tokenizer.pkl # Trained BPE tokenizer
6tokenizer_config.json # Tokenizer configuration
7training_metrics.json # Training metrics
8README.md # This file
1@misc{autoresearch_tinystories_depth8,
2 title={AutoResearch-tinystories-depth8},
3 author={Dustin Loring},
4 year={2026},
5 howpublished={\url{https://huggingface.co/quik-models/curious-thunder-77}}
6}}