A 4B parameter Small Language Model fine-tuned for human-quality customer service in Portuguese. Built with LoRA/QLoRA on Qwen3-4B-Instruct using Unsloth, optimized for consultative conversations, needs discovery, and empathetic support.
Base Model
Qwen3-4B-Instruct-2507 — a lightweight instruction-following model with strong multilingual performance.
Training Dataset
~10,000 synthetic examples generated by GPT-4o-mini following a structured prompt covering:
Consultative conversations (5,000)
Objection handling (2,000)
Memory & context continuity (2,000)
Bad vs ideal responses (1,000)
Categories include indecisive, busy, angry, curious, price inquiries, scheduling, complaints, and more. All examples are in Brazilian Portuguese.
4-bit QLoRA (default) or full LoRA — both supported via --qlora / --lora flags.
Results are merged to 16-bit for inference and GGUF export.
Intended Use
Customer service chatbots in Portuguese
Consultative sales conversations
First-contact qualification
Follow-ups and appointment scheduling
Complaint handling with empathy
Not intended for: general-purpose chat, factual question answering, code generation, or languages other than Portuguese.
Quantization
The uploaded GGUF file is f16 (16-bit float), preserving full fine-tuned quality.
File: slm-customer-service-f16.gguf
How to Use with Ollama
bash
1# Download the GGUF and create a Modelfile:2FROM ./slm-customer-service-f16.gguf
3PARAMETER stop "<|im_end|>"4PARAMETER temperature 0.75PARAMETER top_p 0.96PARAMETER num_predict 20078# Import and run:9ollama create slm-service -f Modelfile
10ollama run slm-service
<|im_start|>system
You are a helpful customer service assistant.<|im_end|>
<|im_start|>user
Olá, gostaria de saber mais sobre o seguro.<|im_end|>
<|im_start|>assistant
Claro! Vou ficar feliz em ajudar. Você já tem uma ideia do tipo de cobertura que procura?<|im_end|>