One of the largest LLMs possible to create in Google Colab.
Trained using a corpus of 70GB of text, nearly twice that of OpenWebText.
Requires 10GB of VRAM.
Technical Information
Layers
36
Heads
20
Embeddings
1280
Context Window
32768 tokens
Tokenizer
GPT-2 BPE
System Tokens
💻🌀
Input Tokens
📋📄
Thinking Tokens
🧠💡
Output Tokens
✅❌
Place the first of each tokens before the text, and the second after the text.