This is a 8-layer decoder-only Transformer trained on the opencaption-finegrained-clone dataset for 0.3 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 1.671475 (perplexity: 3.1854) on the held-out validation set.
Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset.
1Once upon a time, A perfect laptop.
2
3--- a casual weekend scene captures not just in the foreground right of the person, energetic enjoying the spirit of the physical of communal aspect of focus shared dynamic interaction within a young everyday individual performance.
4
5In summary: This description captures a dynamic
1A lonely dragon as 19 on 19 19 .
2
3 central 19 19 19
1The opposite of boy is has dark hair near a wearing glasses, upright. focused is looking’s expression or focused interaction, conveying and shadows. The composition emphasizes and and contrasts sharply with the the the vibrant and the of the vibrant, textured, sense of nostalgia and classic classic.
2
3
1My name is ARVII describ The posture suggests a base, suitcases may be hand-painted.
2
3 eleganceth century, in the background, history, natural and the scale of the and of the location. within a The overall mood is quiet and somewhat quiet, highlighting of a
1import torch
2import pickle
3import json
4from train import GPT, GPTConfig, Tokenizer
5
6# Load config
7with open('config.json', 'r') as f:
8 config_dict = json.load(f)
9config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__})
10
11# Load model
12model = GPT(config)
13state_dict = torch.load('model.pt', map_location='cpu')['state_dict']
14model.load_state_dict(state_dict)
15model.eval()
16
17# Load tokenizer
18with open('tokenizer.pkl', 'rb') as f:
19 tokenizer = pickle.load(f)
20
21# Generate
22prompt = 'Once upon a time, '
23input_ids = tokenizer.encode(prompt)
24x = torch.tensor([input_ids], dtype=torch.long)
25with torch.no_grad():
26 for _ in range(50):
27 logits = model(x)
28 probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1)
29 next_token = torch.multinomial(probs, num_samples=1)
30 input_ids.append(next_token.item())
31 x = torch.tensor([input_ids], dtype=torch.long)
32print(tokenizer.decode(input_ids))
1model.pt # Model weights
2config.json # Model architecture config
3dataset.txt # Dataset name used for training
4token_bytes.pt # Token byte mappings
5tokenizer.pkl # Trained BPE tokenizer
6tokenizer_config.json # Tokenizer configuration
7training_metrics.json # Training metrics
8README.md # This file
1@misc{autoresearch_opencaption-finegrained-clone_depth8,
2 title={AutoResearch-opencaption-finegrained-clone-depth8},
3 author={Dustin Loring},
4 year={2026},
5 howpublished={\url{https://huggingface.co/quik-models/faithful-bee-70}}
6}}