This is a 10-layer decoder-only Transformer trained on the TinyStories (karpathy/tinystories-gpt4-clean) dataset for 1.0 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 0.462273 (perplexity: 1.3777) on the held-out validation set.
Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset.
1Once upon a time, there was a big, red ball. The ball had a friend, a little boy named Tim. Tim liked to play with the ball every day.
2One day, Tim and the ball went to the park. They played with the ball and had lots of fun. They laughed and played all day. The sun was shining, and they were very happy.
3At the end of the day, Tim and the ball were tired. They sat under a big tree and talked. Tim said,
1The opposite of boy is a boy who has a boy who has a boy who is very boy. The boy has a boy who is very kind and he has a boy who is very kind. The boy has a boy who has a boy who is very kind and he is very brave. He says he is a boy who has a boy who is very kind and he is very kind. He says he is a boy who likes to play and he loves his boy very much.
2The boy is very happy that he has
1import torch
2import pickle
3import json
4from train import GPT, GPTConfig, Tokenizer
5
6# Load config
7with open('config.json', 'r') as f:
8 config_dict = json.load(f)
9config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__})
10
11# Load model
12model = GPT(config)
13state_dict = torch.load('model.pt', map_location='cpu')['state_dict']
14model.load_state_dict(state_dict)
15model.eval()
16
17# Load tokenizer
18with open('tokenizer.pkl', 'rb') as f:
19 tokenizer = pickle.load(f)
20
21# Generate
22prompt = 'Once upon a time, '
23input_ids = tokenizer.encode(prompt)
24x = torch.tensor([input_ids], dtype=torch.long)
25with torch.no_grad():
26 for _ in range(50):
27 logits = model(x)
28 probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1)
29 next_token = torch.multinomial(probs, num_samples=1)
30 input_ids.append(next_token.item())
31 x = torch.tensor([input_ids], dtype=torch.long)
32print(tokenizer.decode(input_ids))
1model.pt # Model weights
2config.json # Model architecture config
3dataset.txt # Dataset name used for training
4token_bytes.pt # Token byte mappings
5tokenizer.pkl # Trained BPE tokenizer
6tokenizer_config.json # Tokenizer configuration
7training_metrics.json # Training metrics
8README.md # This file
1@misc{autoresearch_tinystories_depth10,
2 title={AutoResearch-tinystories-depth10},
3 author={Dustin Loring},
4 year={2026},
5 howpublished={\url{https://huggingface.co/quik-models/sleek-sun-138}}
6}}