A simple transformer decoder-only model trained from scratch on the TinyStories dataset.
Model Details
Parameter
Value
Layers
12
Hidden Size
384
Attention Heads
6
Attention Type
MHA
FFN Type
ReLU
Vocab Size
50,257
Max Sequence Length
256
Jounery
this is my first model which I trained from scratch kind of proud of myself for this i know this isn't huge but honest work, it's going to defnitely improve from here