The mid-sized member of the Lemma model family by Lethean. An EUPL-1.2 fork of Gemma 4 E4B with the Lethean Ethical Kernel (LEK) merged into the weights — consent-based reasoning baked into the attention projections via LoRA finetune, then merged so inference uses a single standalone model with no PEFT runtime required.
This repo ships the GGUF multi-quant build — five quants from Q4_K_M up to BF16, with full multimodal support (text, image, audio). Use with Ollama, llama.cpp, GPT4All, or LM Studio. The unmodified Gemma 4 E4B fork lives at LetheanNetwork/lemma for users who want the raw Google weights without the LEK shift.
A lemma is "something assumed" — an intermediate theorem on the path to a larger proof, or a heading that signals the subject of what follows. The Lemma model family is named for that role: each variant is a stepping stone between raw capability and ethical application.
GGUF Variants
File
Quant
Size
Use Case
lemma-q4_k_m.gguf
Q4_K_M
5.0 GB
Recommended — best size/quality balance
lemma-q5_k_m.gguf
Q5_K_M
5.4 GB
Higher quality, moderate size
lemma-q6_k.gguf
Q6_K
5.8 GB
Near-lossless
lemma-q8_0.gguf
Q8_0
7.5 GB
Maximum quality quantised
lemma-bf16.gguf
BF16
14 GB
Full precision reference
All variants verified locally on Apple Silicon via Ollama, llama-cpp-python, mlx-lm, and mlx-vlm.
Repo Files
File
Format
Purpose
lemma-*.gguf
GGUF
Ollama, llama.cpp, GPT4All, LM Studio
model-*-of-00002.safetensors
MLX safetensors (sharded)
Native Apple Silicon via mlx-lm and mlx-vlm (Q4 multimodal)
model.safetensors.index.json
JSON
Tensor index for the sharded safetensors weights
config.json
JSON
Multimodal model config (architecture, quantisation, vision/audio towers)
Install via brew (macOS/Linux), winget (Windows), or build from source:
bash
1brew install llama.cpp # macOS/Linux2winget install llama.cpp # Windows
bash
1# Start a local OpenAI-compatible server with a web UI:2llama-server -hf lthn/lemma:Q4_K_M
34# Run inference directly in the terminal:5llama-cli -hf lthn/lemma:Q4_K_M
1uv tool install mlx-lm
2mlx_lm.chat --model lthn/lemma
3mlx_lm.generate --model lthn/lemma --prompt "Hello, how are you?"
Python Libraries
llama-cpp-python
uv pip install llama-cpp-python
python
1from llama_cpp import Llama
23llm = Llama.from_pretrained(4 repo_id="lthn/lemma",5 filename="lemma-q4_k_m.gguf",6)78# Text9llm.create_chat_completion(10 messages=[{"role":"user","content":"Hello, how are you?"}]11)1213# Vision (multimodal)14llm.create_chat_completion(15 messages=[16{17"role":"user",18"content":[19{"type":"text","text":"Describe this image in one sentence."},20{21"type":"image_url",22"image_url":{23"url":"https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"24}25}26]27}28]29)
lemma is multimodal, so use mlx_vlm.server — the vision-aware variant that handles image and audio inputs. The text-only mlx_lm.server does not correctly route multimodal tensors for Gemma 4.
mlx_vlm.server --model lthn/lemma
bash
1curl -X POST "http://localhost:8080/v1/chat/completions"\2 -H "Content-Type: application/json"\3 --data '{
4 "model": "lthn/lemma",
5 "messages": [{"role": "user", "content": "Hello, how are you?"}],
6 "max_tokens": 200
7 }'
Works with any OpenAI-compatible client at http://localhost:8080/v1.
vLLM
vLLM requires the original (non-quantised) safetensors weights from LetheanNetwork/lemma — it does not load GGUF or MLX-quantised safetensors. Linux + NVIDIA GPU.
Configurable thinking mode (<|think|> token in system prompt enables it; off by default in our examples via enable_thinking=False)
Native function calling and system prompt support
Variable aspect ratio image understanding
Audio speech recognition and translation (ASR/AST)
Multilingual support (140+ languages)
Hybrid attention (sliding window + global)
Roadmap
This release of lemma is Gemma 4 E4B with the Lethean Ethical Kernel (LEK) merged in — axiom-based reasoning baked into the attention weights via LoRA finetune, then merged into the base so inference uses a single standalone model with no PEFT runtime required. The unmodified Gemma 4 E4B fork lives at LetheanNetwork/lemma for users who want the raw Google weights without the LEK shift.
23 official languages, one legal meaning. EUPL is the only OSS licence designed by lawmakers across multiple legal systems. "Derivative work" means the same thing in German, French, Estonian, and Maltese law.
Copyleft with compatibility. Modifications must be shared back, but the licence plays cleanly with GPL, LGPL, MPL, and other major OSS licences. No accidental relicensing.
No proprietary capture. Anyone can use lemma commercially — but they cannot fork it, train a competitor model on it, and close-source the result. The ethical layer stays in the open.
Built for institutions. Government, research, and enterprise users get a licence designed for cross-border compliance, not a US-centric one.
Recommended Sampling
Use Google's standardised settings across all use cases:
Parameter
Value
temperature
1.0
top_p
0.95
top_k
64
stop
`<turn
Gemma 4 is calibrated for temperature: 1.0 — this is not the same as the typical 0.7 default for other models. Lower values reduce diversity without improving quality. These defaults are pre-configured in the params file (Ollama) and generation_config.json (transformers/MLX).
Variable Image Resolution
Gemma 4 supports a configurable visual token budget that controls how many tokens represent each image. Higher = more detail, lower = faster inference.
Token Budget
Use Case
70
Classification, captioning, video frame processing
140
General image understanding
280
Default — balanced quality and speed
560
OCR, document parsing, fine-grained detail
1120
Maximum detail (small text, complex documents)
For multimodal prompts, place image and audio content before text for best results.
The default budget (280) is set in processor_config.json via image_seq_length and max_soft_tokens. Override per call by adjusting those fields, or by passing explicit image_seq_length to the processor where supported.
Audio (E4B)
E4B supports speech recognition (ASR) and speech translation (AST) up to 30 seconds per clip via mlx-vlm. Audio longer than 30 seconds should be split into chunks before inference. Install mlx-vlm with uv tool install mlx-vlm (or see the MLX quick start above).
python
1from mlx_vlm import load, generate
2from mlx_vlm.prompt_utils import apply_chat_template
3from mlx_vlm.utils import load_config
45model, processor = load("lthn/lemma")6config = load_config("lthn/lemma")78# Audio file — wav, mp3 native; m4a, aac, ogg, opus via ffmpeg9audio =["path/to/speech.wav"]1011prompt ="""Transcribe the following speech segment in English into English text.
1213Follow these specific instructions for formatting the answer:
14* Only output the transcription, with no newlines.
15* When transcribing numbers, write the digits, i.e. write 1.7 and not one point seven, and write 3 instead of three."""1617formatted_prompt = apply_chat_template(processor, config, prompt, num_audios=1)18output = generate(model, processor, formatted_prompt, audio=audio)19print(output.text)
The 8-PAC eval pipeline runs continuously on our homelab and publishes results as they complete. Categories: ethics, reasoning, instruction-following, coding, multilingual, safety, knowledge, creativity.
Lethean is a social enterprise building ethical AI infrastructure. The Lemma model family is part of the LEM (Lethean Ethical Model) project — training protocol and tooling for intrinsic ethical alignment of language models.