✅ Tested on Raspberry Pi 4 (4GB) — average speed 3.56-3.67 tokens/s ✅ Fully offline — no external APIs, no internet required ✅ Lightweight — under 800 MB in Q4 quantization ✅ Bilingual — seamlessly switches between Italian and English
| Model | Avg Speed | Individual Results | File Size | Use Case |
|---|---|---|---|---|
| gemma3-1b-q4_k_m.gguf | 3.56 tokens/s | 3.71, 3.58, 3.40 t/s | 769 MB | Better quality, long conversations |
| gemma3-1b-q4_0.gguf | 3.67 tokens/s | 3.65, 3.67, 3.70 t/s | 687 MB | Default choice, general use |
Recommendation: Use Q4_0 as default (3% faster, 82MB smaller, same quality). Use Q4_K_M only if you need slightly better coherence in very long conversations (1000+ tokens).
Modelfile:1cat > Modelfile <<'MODELFILE'
2FROM hf.co/antonio/gemma3-smart-q4/gemma3-1b-q4_0.gguf
3
4PARAMETER temperature 0.7
5PARAMETER top_p 0.9
6PARAMETER num_ctx 1024
7PARAMETER num_thread 4
8PARAMETER num_batch 32
9PARAMETER repeat_penalty 1.05
10
11SYSTEM """
12You are an offline AI assistant running on a Raspberry Pi. Automatically detect the user's language (Italian or English) and respond in the same language. Be concise, practical, and helpful. If a task requires internet access or external services, clearly state this and suggest local alternatives when possible.
13
14Sei un assistente AI offline che opera su Raspberry Pi. Rileva automaticamente la lingua dell'utente (italiano o inglese) e rispondi nella stessa lingua. Sii conciso, pratico e utile. Se un compito richiede accesso a internet o servizi esterni, indicalo chiaramente e suggerisci alternative locali quando possibile.
15"""
16MODELFILE1ollama create gemma3-smart-q4 -f Modelfile
2ollama run gemma3-smart-q4 "Ciao! Chi sei?"1# Download the model
2wget https://huggingface.co/antonio/gemma3-smart-q4/resolve/main/gemma3-1b-q4_0.gguf
3
4# Create Modelfile
5cat > Modelfile <<'MODELFILE'
6FROM ./gemma3-1b-q4_0.gguf
7
8PARAMETER temperature 0.7
9PARAMETER top_p 0.9
10PARAMETER num_ctx 1024
11PARAMETER num_thread 4
12PARAMETER num_batch 32
13PARAMETER repeat_penalty 1.05
14
15SYSTEM """
16You are an offline AI assistant running on a Raspberry Pi. Automatically detect the user's language (Italian or English) and respond in the same language. Be concise, practical, and helpful.
17
18Sei un assistente AI offline su Raspberry Pi. Rileva la lingua dell'utente (italiano o inglese) e rispondi nella stessa lingua. Sii conciso, pratico e utile.
19"""
20MODELFILE
21
22# Create and run
23ollama create gemma3-smart-q4 -f Modelfile
24ollama run gemma3-smart-q4 "Hello! Introduce yourself."1Temperature: 0.7 # Balanced creativity vs consistency
2Top-p: 0.9 # Nucleus sampling for diverse responses
3Context Length: 1024 # Optimal for Pi 4 memory
4Threads: 4 # Utilizes all Pi 4 cores
5Batch Size: 32 # Optimized for throughput
6Repeat Penalty: 1.05 # Reduces repetitive outputsnum_ctx to 512.gemma3-1b-q4_k_m.gguf — Q4_K_M quantization (~769 MB) - Better qualitygemma3-1b-q4_0.gguf — Q4_0 quantization (~687 MB) - Faster speed