Views
No views yet
TinyStories dataset. This collection explores the boundaries of grammatical coherence, narrative depth, and reasoning capacity at micro-scale.| Parameter | Value | Description |
|---|---|---|
model_type | llama | Underlying transformer architecture |
num_hidden_layers | 4 | Number of transformer decoder layers (depth) |
hidden_size | 176 | Hidden dimension size ($d_{model}$) |
intermediate_size | 432 | MLP gate/up projection dimension |
num_attention_heads | 4 | Number of query attention heads |
num_key_value_heads | 1 | Key-value heads (enables 4:1 ratio GQA) |
head_dim | 44 | Vector dimension per attention head |
max_position_embeddings | 320 | Context window size |
vocab_size | 1,536 | Compressed target vocabulary size |
hidden_act | silu | SwiGLU activation function |
tie_word_embeddings | true | Shared input/output embedding representations |
rope_theta | 600.0 | Custom rotary positional embedding base frequency |
attention_bias / mlp_bias | false | Linear layer bias configuration |
bos_token_id / eos_token_id | 0 / 1 | Special token mappings |
d_model is too small—specifically under roughly 1000 dimensions, and most severely at or below 512 dimensions—there is a severe mathematical mismatch between the low rank of the hidden space and the high rank of the target contextual probability distribution of natural language.hidden_size >= 512 is a hard physical threshold. Below 512 dimensions, the model falls off a cliff. Crucially, the paper proved that even stacking the model to an extremely deep 32 or 64 layers cannot compensate for a narrow width.TinyStories corpus. This bypassed the degenerate latent representation trap, allowing our 4-layer, 176-wide model to train stably to completion without saturation.<sink> token (ID 3) is consistently assigned to Position 0 during both pre-training packaging and downstream user inference, I configured a native "post_processor" directly in the tokenizer.json file.TemplateProcessing, the tokenizer automatically prepends the <sink> token, followed by the standard start-of-sequence token <s> (ID 0) to any input string. This eliminates the need for manual prompt modification, ensuring the attention heads always have a dedicated, permanent coordinate at step 0 to dump their unused activation energy.1 "post_processor": {
2 "type": "TemplateProcessing",
3 "single": [
4 {
5 "SpecialToken": {
6 "id": "<sink>",
7 "type_id": 0
8 }
9 },
10 {
11 "SpecialToken": {
12 "id": "<s>",
13 "type_id": 0
14 }
15 },
16 {
17 "Sequence": {
18 "id": "A",
19 "type_id": 0
20 }
21 }
22 ],
23 "pair": [
24 {
25 "SpecialToken": {
26 "id": "<sink>",
27 "type_id": 0
28 }
29 },
30 {
31 "SpecialToken": {
32 "id": "<s>",
33 "type_id": 0
34 }
35 },
36 {
37 "Sequence": {
38 "id": "A",
39 "type_id": 0
40 }
41 },
42 {
43 "SpecialToken": {
44 "id": "</s>",
45 "type_id": 0
46 }
47 },
48 {
49 "Sequence": {
50 "id": "B",
51 "type_id": 1
52 }
53 }
54 ],
55 "special_tokens": {
56 "<s>": {
57 "id": "<s>",
58 "ids": [0],
59 "tokens": ["<s>"]
60 },
61 "</s>": {
62 "id": "</s>",
63 "ids": [1],
64 "tokens": ["</s>"]
65 },
66 "<sink>": {
67 "id": "<sink>",
68 "ids": [3],
69 "tokens": ["<sink>"]
70 }
71 }
72 }tokenizer(prompt) automatically maps the inputs to a <sink> <s> [Prompt] structure. The result is highly stable long-context attention maps, allowing our tiny 1.49M model to write full, 120+ token stories cleanly.<sink>) into our dataset pipeline.<sink> token (ID 3) to Position 0 of every single 320-token block. This gave the attention heads a dedicated, permanent coordinate at step 0 to dump their unused attention energy. The result is stable long-context attention maps, allowing our tiny 1.49M model to write full, 120+ token stories cleanly._once). If a user inputs a prompt and omits the trailing space (e.g., typing "Once upon a time, a little boy named Paco" instead of "Once upon a time, a little boy named Paco "), the tokenizer is forced to split the final word into character-level fragments (like P + ac + o). This pushes the first step of generation completely out-of-distribution, causing the model to output immediate gibberish.once , upon , a ). This matches the natural way punctuation marks (like commas and periods) attach directly to the preceding word without needing extra embedding dimensions to learn spacing rules.tokenizer.json pre-processing pipeline. We replaced the null normalizer block with a regex-based Replace sequence:1 "normalizer": {
2 "type": "Sequence",
3 "normalizers": [
4 {
5 "type": "Replace",
6 "pattern": {
7 "Regex": "(?:\\s)*$"
8 },
9 "content": " "
10 }
11 ]
12 }1Prompt Context: play. When they got to the park, they found the tree and started to climb. One little boy, named Tim, was scared to climb the tree. He said to his friends, "I don't want to climb. I am scared." His friends told him, "It's okay, Tim. You don't
2Generated: play. When they got to the park, they found the tree and started to climb. One little boy, named Tim, was scared to climb the tree. He said to his friends, "I don't want to climb. I am scared." His friends told him, "It's okay, Tim. You don't. You can do it again."
3Tim was sad and said, "I will help you, friends. I will help you." He climbed the tree and got the tree. The tree was happy and said, "Thank you, Tim!" Tim was happy to help his friends. They all worked together to get the tree. The tree was not scared anymore.
4From that day on, Tim and his friends were the best of friends. They played together every day. The tree was not scared anymore. The tree was not scared anymore. It was happy and could help its friends. And they all lived happily ever after.1Prompt Context: ries. The bird ate them all up and felt much better. Then, something unexpected happened. The bird started to grow bigger and bigger! It turned out that the bird was a magic bird! The bird thanked Lily for saving it and gave her a wish. Lily wished for
2Generated: ries. The bird ate them all up and felt much better. Then, something unexpected happened. The bird started to grow bigger and bigger! It turned out that the bird was a magic bird! The bird thanked Lily for saving it and gave her a wish. Lily wished forgive the bird and they all became friends. From that day on, the bird and Lily were the best of friends.1Prompt Context: . One day, the jeep saw an angry dog. The dog did not want to play. The jeep was sad. The jeep went to find its friends. They all played and had fun. The angry dog saw them and wanted to play too. They all played together and
2Generated: . One day, the jeep saw an angry dog. The dog did not want to play. The jeep was sad. The jeep went to find its friends. They all played and had fun. The angry dog saw them and wanted to play too. They all played together and, and they had lots of fun.temp=0.35, min_p=0.10)Once upon a time, there was a little girl named Lily. She loved to play with her toys and have fun. One day, she found a big box in her room. She was very happy and wanted to see what was inside. Lily opened the box and found a pretty dress. She put on the dress and went outside to play. She saw her friend, Tom, and said, "Look, Tom! I found a dress!" Tom looked at the dress and smiled. He said, "Wow! That's a nice dress!" Lily and Tom played with the dress all day. They took turns wearing it and pretending to be kings and queens. They had so much fun together. At the end of the day, they put the dress back in the box and said, "We had a great day!"
Once upon a time, there was a little boy named Tim. Tim loved to play with his toy car. One day, he found a big box in his room. He was very curious about the box. Tim opened the box and found a toy car. The car was big and red. Tim was very happy. He wanted to play with the car. But when he triedto pull the car, it did not move. Tim was sad. Then, Tim had an idea. He took the car to his mom. She said, "Let's pull the car out of the box." They pushed and pulled. The car started to move! It was not a car at all. It was a magic car! The car could make the car go fast again. Tim and his mom were very happy. They played with the magic car all day.
Once upon a time, a little boy named Paco was very excited. He wanted to go to the park with his mom. He put on his shoes and ran outside. When Paco got to the park, he saw a big slide. He ran up to it and started to slide down. He slid down fast and laughed as he went faster and faster. He felt so happy and excited. When he got to the park, he saw a big, green frog. The frog was very friendly and said, "Hi, Paco! Do you want to slide with me?" Paco smiled and said, "Yes, let's slide together!" So, Paco and the frog slid down the slide together, laughing and having fun.
Once upon a time, a little boy named Paco was playing with his toy car. He was very excited to see what he could do. He ran to his mom and said, "Mom, I want to play with my car!" His mom smiled and said, "Okay, Paco and your car, but be careful." Paco and his mom played with the car, making it go fast. They had so much fun with the car. But then, something unexpected happened. The car's car started to move! It was not a car, but a big turtle named Tina. Tina was very surprised! "Hello, Paco and Mom!" said Tina. Paco was so surprised that he dropped the car. The turtle said, "Thank you for finding my car! I was stuck as a turtle! Let's play together again!" Paco and Tina were happy to have a new friend.
Once upon a time, a little boy named Paco went to the beach. He loved to play in the sand and swim in the water. One day, he saw a big, red ball in the sand. He wanted to play with it, so he ran to get it. As Paco got close to the ball, he heard a voice. "Hey, that's my ball!" said the voice. Paco looked around and saw a little girl named Lily. She was holding the ball in her hand. "Hi, Lily!" said Paco. "I found this ball. It was my favorite toy." Lily smiled and said, "Thank you, Paco! I found it in the sand. I found it in the sand." Paco and Lily played with the ball together, and they became good friends.
Once upon a time, a little boy named Paco went to the beach. He was very excited to play in the sand and swim in the water. He saw a big crab and wanted to play with it. Paco was very excited and jumped in the water. The crab jumped in the water and started to swim with Paco. Paco was having so much fun. But then, something unexpected happened. Paco was not a crab at all! It was a big, friendly crab who lived in the ocean. The crab was not a crab anymore. He was a real crab who lived in the ocean. Pacan was so happy to have a new friend, and he played with the crab all day long.
[None] (Note: The 4.0M model cannot initiate generation from an empty string without a custom prompt handler, as its tokenization structures require a text prefix to align the special token masks).
Once upon a time, there was a little girl named Sue. Sue loved to play with her toys and make music. One day, she found a big box in her room. The box was very old and had many colors. Sue wanted to see what was inside. Sue opened the box and saw a lot of colors. She thought it would be fun to play with the box. So, she took the box and started to drawon the floor. She drew a big sun, a house, and a happy family. Then, she drew a big sun, a green leaf, and a green leaf. Sue was very happy with her new toy. She played with it all day long. But then, she heard a loud noise. It was a big bear! The bear was hungry and wanted to eat Sue. Sue was scared, but she knew she had to be brave. She ran to the bear and said, "Don't worry, Mr. Bear.I will help you find your toy." The bear was very happy and thanked Sue for her help. They played with the toy and had lots of fun. Sue learned that it is good to help others when they need it. And they all lived happily ever after.
TinyStoriesV2) is conceptually designed around simple, child-like narratives, this model has not undergone safety tuning or toxic content filtering. It can output unpredictable, nonsensical, or potentially inappropriate text. Consequently, under no circumstances should this model or its outputs be deemed safe, verified, or appropriate for children or general public interaction.transformers library with zero custom configurations or remote execution flags.1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model_name = "MultivexAI/Aurelius-Llama-v2.0-1.5M-Nano"
4
5tokenizer = AutoTokenizer.from_pretrained(model_name)
6model = AutoModelForCausalLM.from_pretrained(model_name)
7
8# Note: You can omit the trailing space; the tokenizer normalizer will automatically handle it!
9prompt = "Once upon a time, a little boy named Paco"
10inputs = tokenizer(prompt, return_tensors="pt")
11
12outputs = model.generate(
13 **inputs,
14 max_new_tokens=128,
15 temperature=0.60,
16 min_p=0.15,
17 do_sample=True,
18 pad_token_id=tokenizer.eos_token_id,
19 eos_token_id=tokenizer.eos_token_id
20)
21print(tokenizer.decode(outputs[0], skip_special_tokens=True))