FunctionGemma 270M IT — Prepaid Cards Tool-Calling (v2, SafeTensors)
Model description
A fine-tuned version of google/functiongemma-270m-it
(Gemma 3 270M, 268M params) that recognizes prepaid-card intents in chat and
emits the correct tool call:
Buy a Digital Prepaid Visa or Virtual Prepaid Mastercard
get_card_balance(card_number)
Check the balance of a card
get_transaction_history(card_number, limit?)
List a card's transactions
Trained on the v2 dataset: 107 languages, multi-turn conversations (card
number in one message, request in another; clarification loops; full
call→response loops), and realistic user noise (typos, text-speak, dropped
articles, scrambled word order) so the model works with how people actually
type.
Intended uses & limitations
Intended uses
Chat agents that buy prepaid cards, answer balance questions, and show
transaction history, in many languages and with noisy/multi-turn input.
Distillation target: a small model that a backend can drive via the
standard FunctionGemma <start_function_call>… protocol.
Limitations & biases
Synthetic training data. All conversations are generated from
hand-written templates; the model has not seen real user traffic.
Uneven language quality. English and ~30 major languages are the most
richly covered; the 20+ low-resource languages were translated by hand and
contain approximations. Held-out-language accuracy (89.8% in v1) lags
English slightly.
No backend. The model only emits tool calls; it cannot check balances
or buy cards itself.
Security note: like all small models it can mis-parse card numbers
under heavy noise — validate tool arguments before executing payments.
Gemma license applies (base model license).
How to use
python
1from transformers import AutoTokenizer, AutoModelForCausalLM
2import torch, json
3from transformers.utils import get_json_schema
45model = AutoModelForCausalLM.from_pretrained(6"Qrzysztof/functiongemma-270m-it-prepaid-cards-v2",7 dtype=torch.bfloat16, attn_implementation="eager")8tokenizer = AutoTokenizer.from_pretrained("Qrzysztof/functiongemma-270m-it-prepaid-cards-v2")910defpurchase_card(amount:float, card_type:str, email:str="", currency:str="USD")->str:...11defget_card_balance(card_number:str)->str:...12TOOLS =[get_json_schema(purchase_card), get_json_schema(get_card_balance)]1314messages =[15{"role":"developer","content":"You are a model that can do function calling with the following functions"},16{"role":"user","content":"i wanna buy a 20 dollar card plz"},# noisy input works17]18inputs = tokenizer.apply_chat_template(messages, tools=TOOLS, add_generation_prompt=True,19 return_dict=True, return_tensors="pt")20out = model.generate(**inputs, max_new_tokens=128)21print(tokenizer.decode(out[0][len(inputs["input_ids"][0]):], skip_special_tokens=False))22# <start_function_call>call:purchase_card{"amount": 20, "card_type": "digital_prepaid_visa", ...}<end_function_call>
Training details
Parameter
Value
Base model
google/functiongemma-270m-it (Gemma 3 270M)
Method
Full fine-tune (all 268M params), TRL SFTTrainer
Data
Qrzysztof/functiongemma-prepaid-cards-tool-calling-v2 — ~1,800 samples/epoch (balanced across 107 languages & intents)
Epochs
3
Batch
8 (T4, bf16, eager attention)
Max length
1024
LR / schedule
5e-5, constant, 50 warmup steps
Hardware
Google Colab T4 GPU
Per-epoch checkpoints: checkpoint/epoch-{1,2,3}.
Evaluation
Method: greedy decoding over the held-out v2 test split (never in training;
5 languages fully held out — ja, ko, ar, sw, ur). A sample counts as
correct when the generated text contains the expected tool name and no other
tool name (for text-response samples: when it contains no tool call).
Bucket
v1
v2
Overall (295 samples)
91.5% (v1 split)
89.5% (harder v2 split)
purchase_card
92.9%
90.3%
get_card_balance
95.7%
87.0%
get_transaction_history
80.9%
83.8%
Multi-turn chains
98%
100%
Seen languages
94.0%
94.8%
Held-out languages
89.8%
86.0%
Cross-format comparison (subset): torch / GGUF Q8_0 / MLX 8-bit all score
40/40 (100%) on the same 40 prompts; ONNX: see the
ONNX repo.
Fine-tuning from this model
This model was fine-tuned with the tutorial below; you can use it as the starting point for a new tool set (or fine-tune google/functiongemma-270m-it directly).
Fine-tuning tutorial
A complete, minimal fine-tune of a FunctionGemma-class model on this data
(follows the official
FunctionGemma fine-tuning guide).
1. Setup
bash
1pip install torch transformers trl datasets accelerate
2huggingface-cli login # accept the gemma license for google/functiongemma-270m-it
2. Load the dataset and normalize messages
The Hub dataset stores messages/tools as JSON strings (Arrow cannot infer
the nested schema), and TRL's SFTTrainer needs a uniform struct schema, so
normalize first:
python
1import json
2from datasets import load_dataset
3from transformers import AutoModelForCausalLM, AutoTokenizer
45defnormalize_messages(msgs):6 out =[]7for m in msgs:8 n ={"role": m["role"],"content": m.get("content")or"","name":None,9"tool_call_id": m.get("tool_call_id"),"tool_calls":None}10if m["role"]=="tool":11 n["name"]= m["content"]["name"]12 n["content"]= json.dumps(m["content"]["response"], ensure_ascii=False)13if m.get("tool_calls"):14 n["tool_calls"]=[{"id": tc.get("id"),"type": tc.get("type","function"),15"function":{"name": tc["function"]["name"],16"arguments": json.dumps(tc["function"]["arguments"], ensure_ascii=False)}}17for tc in m["tool_calls"]]18 out.append(n)19return out
2021defrows_to_dataset(rows):22from datasets import Dataset
23return Dataset.from_list([{24"messages": normalize_messages(r["messages"]),25"tools": json.dumps(r["tools"], ensure_ascii=False),26}for r in rows])2728ds = load_dataset("Qrzysztof/ecommerce-chat-tool-calling", token=HF_TOKEN)["train"]29train_rows =[{"messages": json.loads(r["messages_json"]),"tools": json.loads(r["tools_json"])}30for r in ds if r["split"]=="train"]31train_ds = rows_to_dataset(train_rows)
TRL applies the FunctionGemma chat template with the per-sample tools
column; assistant_only_loss=True (default) masks everything but the model's
own turns, so it learns to emit tool calls — not to copy the schema.
4. Evaluate (greedy success rate)
python
1ok =02for item in test_rows:3 inputs = tokenizer.apply_chat_template(item["messages"][:-1], tools=item["tools"],4 add_generation_prompt=True, return_tensors="pt")5 out = model.generate(**inputs, max_new_tokens=256)6 output = tokenizer.decode(out[0][len(inputs["input_ids"][0]):], skip_special_tokens=False)7 expected =<expected tool name / args from expected_json>8 ok += expected-tool-in-output and no-other-tool-in-output
Keep noise digit-safe: never corrupt the values the model must extract
(prices, ids). The noise.py engine skips any token containing digits.
Use deterministic train/test splits (by template_id) and hold out
whole languages + (for the e-commerce set) whole schemas — that is the
only honest way to measure generalization.
Balance the training subset per (language, intent) — cap the big
buckets instead of letting English dominate.
Training
packing=False for tool-calling data; packed sequences splice mid-call.
max_length ≥ longest sample + a margin; ~1024 covers these datasets.
Constant LR + short warmup (the official guide's defaults) work well.
Upload a checkpoint to the Hub after every epoch — Colab VMs die
mid-run, and the last good epoch is always recoverable.
Evaluation
Always evaluate with greedy decoding for comparability across formats
and runs.
Score two things separately: tool-name selection and argument fidelity
(query + every filter key:value pair).
Compare every exported format (SafeTensors / GGUF / MLX / ONNX) on the
same prompts — quantization changes results.
Deployment
Validate tool arguments server-side before executing anything (a small
model can garble a card number under heavy noise).
In a live agent, follow the FunctionGemma full loop: model call → backend
executes → tool response → model continues; never let the model see or
emit secrets.
For browser deployment use the fp16 ONNX file; for low-end hardware the
Q8_0 GGUF or MLX 8-bit; for exact reference behavior the SafeTensors model.