This repository contains a GGUF quantized version of microsoft/Phi-3.5-mini-instruct, prepared for use with llama.cpp and compatible runtimes, and used as the core instruct model inside the meeTARA empathetic assistant.
Quantization: Q4_K_M (good quality vs RAM / speed)
Intended use: Standalone intelligent assistant with baked-in domain detection, emotional intelligence, and structured responses for local / offline inference.
✨ Standalone Intelligence: This GGUF model includes 20 layers of intelligence baked directly into the chat template. No backend code required - download and use with llama.cpp, Ollama, or any GGUF-compatible runtime.
Available files
Filename
Quant type
Size
Notes
meetara-phi-3.5-mini-instruct-gguf-Q4_K_M.gguf
Q4_K_M
~2.2G
Default quant, recommended
More quantizations (e.g., Q5_K_M, Q8_0) can be added later to this repo as additional .gguf files.
Prompt format (recommended)
The model uses a Qwen‑style chat template. A simple, robust pattern is:
<|im_start|>system
You are meeTARA, an emotionally intelligent AI assistant built on top of a Phi-3.5‑3.8B‑Instruct base model. Always answer clearly, kindly, and with practical steps the user can take.
<|im_end|>
<|im_start|>user
{user_message}
<|im_end|>
<|im_start|>assistant
Example:
<|im_start|>system
You are meeTARA, an emotionally intelligent AI assistant built on top of a Phi-3.5‑3.8B‑Instruct base model. Always answer clearly, kindly, and with practical steps the user can take.
<|im_end|>
<|im_start|>user
How can I improve my sleep quality and manage stress naturally?
<|im_end|>
<|im_start|>assistant
1./llama-cli \2 -m /path/to/meetara-phi-3.5-mini-instruct-gguf-Q4_K_M.gguf \3 -p "<|im_start|>system You are meeTARA, an emotionally intelligent AI assistant built on top of a Phi-3.5‑3.8B‑Instruct base model. Always answer clearly, kindly, and with practical steps the user can take. <|im_end|> <|im_start|>user How can I improve my sleep quality and manage stress naturally? <|im_end|> <|im_start|>assistant"
Adjust flags like -n (max tokens), --temperature, --top_p, --top_k, etc. according to your hardware and latency/quality trade‑offs.
This GGUF model includes baked-in intelligence that works without any backend code. The model automatically detects domains, emotions, intent, and context through a 20-layer detection system:
Intelligence Layers
Layer
Feature
Description
1
🚨 Refusal Patterns
Safety-first harmful request detection
2
🧩 Contextual Patterns
Multi-word phrase disambiguation (python code vs snake)
3
📊 N-gram Patterns
Bigram/trigram detection for better context
4
🔗 Semantic Clusters
Related keyword groups boost domain confidence
5
👤 Entity Patterns
Personal context, time-sensitive, beginner/expert
6
🎯 Intent Signals
What user wants: learn, fix, decide, create, validate
Layer 16: Contextual structure with empathetic opening (e.g. 2–5 sections when helpful)
Expected Response: Empathetic opening, clear answer and key details, practical steps, and safety disclaimer. Structure adapts to question complexity (simpler questions get a more direct answer).
Example 2: Technology Domain with Context Awareness