Fine-tune a small open LLM (Qwen2.5-Coder-1.5B, ≤ 3B) to turn natural-language questions into
executable SQLite queries — the engine for a fintech chatbot that lets non-technical teams pull
data without writing SQL.
The whole project is built around free resources: training on a Google Colab T4, inference on
an ordinary laptop CPU via a quantized GGUF served through Ollama — no GPU required to run
it. A LoRA adapter (~1% of parameters trained) sits on top of the frozen base model and is exported
to a ~1 GB GGUF for local use.
What it does: given a database schema (DDL) + a question, it returns one runnable SQLite query.
What it's for: a data-access chatbot — plain English in, executable SQL out.
Result of the fine-tune
Evaluated on 200 held-out test examples, greedy decoding, identical prompts for both models.
The numbers below are the final max_new_tokens=512 run (Result/*_preds_512.json).
Model
Valid rate
Executable rate
Baseline (no fine-tune)
99.0%
42.5%
Fine-tuned (LoRA, 512-tok)
99.5%
75.5%
Gold queries (ceiling)
100.0%
99.5%
Fine-tuning raised the executable-query rate from 42.5% → 75.5% (+33 pp absolute, +78%
relative), recovering roughly half the gap to the gold ceiling.
Ollama downloads the GGUF for you, no manual steps:
ollama run hf.co/SkibidiBreaddd/qwen2.5-coder-1.5b-t2sql-gguf
Recommended — build with the baked-in system prompt
This applies the Text2SQL system prompt and temperature 0 from report/Modelfile,
so you get deterministic, prompt-correct output:
bash
1# 1. download just the GGUF (~1 GB)2huggingface-cli download SkibidiBreaddd/qwen2.5-coder-1.5b-t2sql-gguf \3 --include "*.gguf" --local-dir ./gguf
45# 2. build a local Ollama model from the Modelfile6ollama create t2sql -f report/Modelfile
78# 3. run it9ollama run t2sql
You then also get an OpenAI-compatible HTTP API on localhost:11434. Prompt it with the schema DDL
followed by the question (same order used in training).