A 62M-parameter GPT trained completely from scratch on a single 8GB-RAM NVIDIA Jetson device, without cloud infrastructure or multi-GPU setups. Base pretrained checkpoint for raw text completion, not instruction following.
Overview
G0-nano-base is a small decoder-only causal language model trained end-to-end under an 8GB unified-memory constraint. The project focuses on making the full training process — tokenizer, pretraining, fine-tuning infrastructure and export — work on modest hardware.
This is the base checkpoint. It predicts the next token and completes text; it is not a chat model and should not be expected to follow instructions.
Model variants
The instruction-tuned version of the same model is available as G0-nano-instruct.
What this version adds
This checkpoint is the pretrained foundation of the G0 Nano model line. It does not include supervised instruction fine-tuning or a chat format.
Architecture
Llama-style decoder-only Transformer:
Property
Value
Parameters
62.1M, with embeddings shared with the language-model head
Pretraining data: approximately 1.5B tokens of English web and book text
Sources: FineWeb-Edu, BookCorpus, OpenWebText, PG-19 and WikiHow
Objective: causal next-token prediction
Training hardware: a single NVIDIA Jetson with 8GB of unified memory
Usage
Hugging Face Transformers
python
1from transformers import AutoModelForCausalLM, AutoTokenizer
23model_id ="AZERDSQ/G0-nano-base"4tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)5model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True)67inputs = tokenizer("The city of Paris is", return_tensors="pt")8outputs = model.generate(9**inputs,10 max_new_tokens=50,11 do_sample=True,12 top_k=50,13 temperature=0.8,14)15print(tokenizer.decode(outputs[0]))
trust_remote_code=True is required because this repository uses a custom Transformer implementation.
Ollama
ollama run azerdsq/g0-nano-base "The city of Paris is"
This is a base model: it completes text rather than answering questions.
Limitations
62M parameters impose a hard limit on factual knowledge; expect fluent but frequently incorrect completions on knowledge-intensive prompts.
Maximum context length is 1024 tokens.
English-only training data.
Single-sequence generation only; padded batched inference is not supported by the custom model code.
No instruction tuning and no chat format.
This model should not be used for high-stakes decisions, factual verification, medical advice, legal advice or autonomous actions.
License
Apache 2.0. This release contains model weights and the code required to load them; it does not include the training data or private training infrastructure.