A QLoRA adapter for google/gemma-4-E2B-it (5.1B raw / ~2B effective parameters, MatFormer) specialized in e-commerce catalog tasks with Magento 2 conventions: attribute extraction to JSON, product Q&A, search-relevance classification, and product ranking. Fourth model in a same-data, same-recipe, same-harness comparison with qwen3.5-4b-ec-magento, ministral3-3b-ec-magento, and phi4-3b-ec-magento.
The adapter was trained on unsloth/gemma-4-E2B-it-unsloth-bnb-4bit (unsloth's dynamic 4-bit export — no plain NF4 exists for this family); the GGUF is the adapter merged into the bf16 weights.
Path
What
Size
For
/ (root)
LoRA adapter (PEFT)
80MB
transformers/peft/unsloth on top of the base model
Trained on 56,253 instruction samples: 70% ECInstruct (generic e-commerce) + 30% synthetic Magento-schema data generated from the Magento Luma sample catalog (products fully disjoint between train and eval). Four task shapes:
Attribute extraction — product text (or a raw Magento custom_attributes payload) → JSON:
Absent attributes are reported as "None" rather than hallucinated.
Product QA — a question answered strictly from given product data.
Relevance classification — query + product → graded relevance option (ESCI-style A–D).
Relevance ranking — query + lettered product list → ranked letters (B,A,C).
Usage — adapter (unsloth / peft)
Needs ~10GB free GPU memory for a straightforward load (the per-layer-embeddings
table and audio/vision towers are unquantized). On smaller GPUs see the
8GB gotchas below.
python
1from unsloth import FastLanguageModel
23model, tokenizer = FastLanguageModel.from_pretrained(4"gabrielgts/gemma4-e2b-ec-magento", max_seq_length=2048, load_in_4bit=True)5FastLanguageModel.for_inference(model)67# No system message — Gemma has no separate system role and the model was8# trained without one. Greedy decoding recommended.9messages =[{"role":"user","content":10"Extract the value of the target attribute from the given product information "11"and output it as JSON. If the attribute is not present, output None as the value.\n\n"12"target attribute: size\nproduct title: Puma Suede green sneakers size 43"}]13text = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)14inputs = tokenizer(text=text, return_tensors="pt").to("cuda")15out = model.generate(**inputs, max_new_tokens=64, do_sample=False)16print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))17# [{"attribute": "size", "value": "43"}]
This model was trained on an 8GB RTX 3070, which required CPU-offloading the
frozen 4.4GB per-layer-embeddings table and the audio/vision towers:
Set llm_int8_enable_fp32_cpu_offload: trueinside the checkpoint'sconfig.json quantization_config (transformers 5.5 ignores the flag when
passed via from_pretrained — BitsAndBytesConfig lost
get_loading_attributes, so user flags never merge for pre-quantized repos).
Pass a fully disjointdevice_map (explicit per-child entries; maps with
a "" catch-all plus nested overrides are silently ignored by the 5.5 loader).
Loading this adapter on top of an offloaded base: avoid
PeftModel.from_pretrained (it re-dispatches and un-offloads the model);
use peft.inject_adapter_in_model + set_peft_model_state_dict instead.
None of this applies on GPUs with ≥10GB free memory.
Training recipe
Parameter
Value
Method
QLoRA (4-bit dynamic-quant base, bf16 compute) via Unsloth
LoRA
r=8, alpha=16, targets q/k/v/o/gate/up/down_proj (language model only; PLE, audio and vision towers frozen)
Greedy decoding, identical prompts across all models; base models evaluated zero-shot with the same harness. Quantization parity caveat: this family trained on unsloth's dynamic 4-bit quant (selected layers kept at higher precision) while the three siblings used plain NF4 — no plain NF4 export exists for gemma4. Direction of bias: slightly favors this model.
Magento held-out set (2,969 samples, 475 products never seen in training):
Task · metric
Base
this model
ministral3-3b
phi4-mini
qwen3.5-4b
Attribute extraction · F1
0.081
0.933
0.945
0.946
0.938
Product QA · token-F1
0.043
0.953
0.954
0.950
0.943
Relevance classification · accuracy
0.311
0.961
0.972
0.959
0.962
Relevance rank · top-1
0.418
0.776
0.793
0.783
0.799
ECInstruct held-out set (2,000 samples):
Task · metric
Base
this model
ministral3-3b
phi4-mini
qwen3.5-4b
Attribute extraction · F1
0.000
0.641
0.654
0.616
0.646
Query→product rank · top-1
0.015
0.647
0.632
0.645
0.650
Relevance classification · accuracy
0.000
0.662
0.655
0.650
0.685
Answerability · accuracy
0.465
0.728
0.735
0.733
0.780
With ~2B effective parameters, this model stays within ~2 points of siblings nearly twice its effective size on every metric — the best capability-per-effective-param in the comparison.
Limitations — read before relying on the numbers
The Magento eval is synthetic-on-synthetic. Eval tasks were generated with the same templates as training data (products fully disjoint). It validly measures schema adherence — JSON format, Magento attribute vocabularies, the None-when-absent rule — but overstates production quality on real catalogs and real user queries.
English only; fine-tuned on structured data — expect degraded general chat and multilingual ability vs the base (drop the adapter to recover). The audio/vision towers are untouched, but this repo's GGUF is text-only.
Use greedy decoding (do_sample=False / temperature 0) — that's how it was evaluated.