An experimental fine-tune of
Qwen/Qwen3.5-4B
for Korean-capable function calling, agent-oriented tool use, and short
coding tasks on Apple Silicon using MLX.
This is an independent personal research and portfolio project. It is not
affiliated with or endorsed by Qwen, Alibaba Cloud, Nous Research, Glaive AI,
OpenAI, Apple, MLX, Hermes Agent, or OpenClaw.
Available formats
This repository publishes the same fine-tuned checkpoint (promotion step
3,900) in two formats:
Hugging Face safetensors, bf16 text causal LM (Qwen3_5ForCausalLM)
transformers (CPU or CUDA GPU), vLLM
GPU/CPU inference without Apple Silicon; full-precision weights
The hf-bf16/ folder is the bf16 (non-quantized) form of the same fused
step-3900 checkpoint, exported for non-MLX runtimes. Notes:
trust_remote_code is not required with transformers >= 5.12 (native
qwen3_5 support).
hf-bf16/ is text-only. The base model is a vision-language
architecture, but this fine-tune contains only the language-model weights;
image input is not supported.
CPU inference with transformers is numerically correct but slow: the
gated-delta linear-attention fast-path kernels (causal-conv1d,
flash-linear-attention) are CUDA-only, so CPU falls back to a slower torch
implementation. Use a GPU for throughput.
Formats may be updated independently after future promotions; pin an exact
commit for reproducibility.
Model summary
Base model:Qwen/Qwen3.5-4B
Training: multi-stage LoRA supervised fine-tuning
Adapter size: approximately 12.4 million trainable parameters
Published artifact: fused model, quantized to 6 bits
Quantization: 6-bit, group size 64
Runtime: MLX / MLX-LM
Languages evaluated: Korean and English
Primary tasks: function calling, tool selection, abstention, and short coding
Status: experimental
The published model is a single fused and quantized artifact. A separate
LoRA adapter is not required.
About the repository name
The repository retains the historical distilled name.
Technically, this model was trained through multi-stage LoRA supervised
fine-tuning on a mixture of public synthetic instruction datasets,
function-calling datasets, and maintainer-generated examples. No claim of
logit-level knowledge distillation is made.
Some public datasets used during training contain model-generated
instruction or response data. This may be understood as indirect behavioral
distillation through supervised fine-tuning, rather than direct
teacher-student logit distillation.
Training stages
Stage 1 — function calling
approximately 7,408 Hermes-formatted function-calling examples;
one training epoch;
optimized primarily for English tool selection and structured calls.
Stage 2 — Korean and abstention continued SFT
continued from the Stage 1 adapter;
added Korean instructions and examples where the model should answer
directly rather than invoke a tool;
improved Korean response compliance and tool-call abstention.
Stage 3 — rebalanced agentic and coding SFT
The current training mixture contains approximately 19,697 records,
including:
Korean direct-answer and abstention examples;
Korean instruction examples;
maintainer-generated Korean tool-calling examples;
prior continued-SFT examples;
Hermes-formatted function-calling examples;
Glaive function-calling examples;
Magicoder coding examples.
The mixture was rebalanced to reduce unconditional tool calling. The
approximate tool-call to no-tool-call ratio is 1.4:1.
Dataset transformations, duplication factors, filtering, and formatting may
change the effective sampling distribution. Record counts should therefore
not be interpreted as counts of unique source examples.
Data provenance
The training mixture includes publicly available datasets and
maintainer-generated examples.
No explicit license declaration identified in the dataset card; refer to the source repository and underlying content terms
5,500 source samples before weighting
Korean instruction and direct-answer/abstention construction
Korean tool calls
Maintainer-generated
Maintainer-authored
Weighted by 4×
Synthetic or manually curated Korean tool-use examples
Continued-SFT examples
Maintainer-curated
Derived from the sources listed above and maintainer-authored examples
Weighted by 3×
Examples retained from earlier training stages
Dataset record counts describe the effective training mixture after
sampling, filtering, transformation, and duplication. They do not represent
the number of unique source records.
The term Hermes-compatible describes the conversation serialization and
tool-call format. It does not imply that every training example originated
from Nous Research.
The Hermes function-calling and Glaive function-calling repositories identify
their datasets as Apache-2.0 at the time of this release.
The Magicoder OSS-Instruct dataset contains synthetic coding instructions
generated from open-source code snippets. Individual source-code snippets
and repositories may be subject to their own licenses. This project does not
relicense upstream source code or dataset content.
The KoAlpaca v1.1a dataset contains Korean question-and-answer records and
references to original source URLs. Its Hugging Face dataset card does not
provide an explicit repository-level license declaration. Accordingly, this
model card does not characterize KoAlpaca v1.1a as Apache-2.0 or otherwise
permissively licensed. Users should review the KoAlpaca repository and the
applicable terms of the underlying source content independently.
All third-party datasets remain subject to their original licenses, terms,
attribution requirements, provenance limitations, and usage restrictions.
The Apache-2.0 license applied to this model repository does not replace or
override those separate terms.
No private company data, customer data, credentials, or intentionally
collected personally identifiable information was knowingly included in the
published training mixture. If such material is later identified, the
affected release should be reviewed and, where appropriate, withdrawn and
rebuilt.
Training implementation
Training was performed on a single Apple M3 Ultra system with 256 GB of
unified memory.
The MLX training implementation includes:
a custom differentiable chunked gated-delta implementation for the hybrid
linear-attention architecture;
gradient checkpointing;
chunked cross-entropy over a vocabulary of approximately 248,320 tokens;
state carry with truncated backpropagation through time;
PoSE-based positional augmentation;
LoRA over selected state-space and full-attention projections.
These implementation details describe the training system. They are not a
guarantee of long-context accuracy or generalization.
Automated checkpoint promotion
Candidate checkpoints are evaluated every 100 training steps.
The automated promotion system considers:
held-out cross-entropy;
tool-selection probes;
tool-call abstention probes;
Korean response compliance;
conversational parroting checks.
A candidate is not promoted when specified minimum thresholds are violated.
When a checkpoint exceeds the current promotion score, the adapter may be
fused, quantized, validated for serving, and uploaded to this repository.
The promotion score is an internal checkpoint-selection heuristic. It is not
a public benchmark and should not be interpreted as a comprehensive measure
of model quality.
The main revision may change after a successful automated promotion.
For reproducible use, pin an exact Hugging Face commit or immutable version
tag.
Evaluation
The currently reported checkpoint was promoted at training step 3,900.
Metric
Result
Held-out cross-entropy, four-domain internal set
0.7241
Tool-selection score
1.00
Tool-call abstention score
0.80
Korean response ratio
1.00
Conversational parroting
0
The held-out cross-entropy set contains 120 examples, with approximately 30
examples from each of four domains:
Glaive-style function calling;
Magicoder-style coding;
KoAlpaca-style Korean instructions;
Hermes-style function calling.
The tool probe used during promotion contains 14 tool-call cases and 15
abstention cases.
These are small, maintainer-defined internal evaluations. The results have
limited statistical power and may not generalize to other prompts, tools,
languages, coding tasks, or agent frameworks.
A reported score of 1.00 means that all examples in that specific small
evaluation subset passed the implemented scoring rule. It does not imply
perfect tool-selection performance in general.
The held-out set used repeatedly for checkpoint promotion is a validation
set rather than a fully independent final test set. A separate untouched
test set is required for an unbiased final performance estimate.
Context length
The model retains the nominal context configuration inherited from the base
model, potentially up to 262,144 tokens.
The full context range has not been comprehensively validated. Effective
tool-use, retrieval, coding, and Korean-language quality may degrade well
before the configured maximum.
Available context also depends on unified memory, MLX-LM version, KV-cache
configuration, prompt structure, and serving concurrency.
Do not interpret the nominal context configuration as a guarantee of
reliable 256K-token performance.
Function calling
The model was trained using Hermes/Qwen-compatible tool-call formatting and
is intended to work with an OpenAI-style local serving API.
Structured tool_calls output depends on:
the model's chat template;
the MLX-LM version;
the serving parser;
post-quantization template fixes;
client and agent-framework behavior.
The model may produce malformed JSON, wrong tool names, missing arguments,
fabricated values, unsafe arguments, or plain text instead of a tool call.
“OpenAI-compatible” refers only to an API convention and does not imply
affiliation with OpenAI or complete compatibility with every OpenAI client.
Usage
Install MLX-LM:
pip install --upgrade mlx-lm
Generate text:
bash
1python -m mlx_lm generate \2 --model idah4/qwen3.5-4b-distilled \3 --prompt "Write a Python function to merge two dictionaries."
Run a local OpenAI-style server:
bash
1python -m mlx_lm server \2 --model idah4/qwen3.5-4b-distilled
Hugging Face transformers (hf-bf16/, CPU or CUDA)
python
1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
34repo ="idah4/qwen3.5-4b-distilled"5tok = AutoTokenizer.from_pretrained(repo, subfolder="hf-bf16")6model = AutoModelForCausalLM.from_pretrained(7 repo, subfolder="hf-bf16",8 dtype=torch.bfloat16,9 low_cpu_mem_usage=True,10)11messages =[{"role":"user","content":"Write a Python function to merge two dictionaries."}]12text = tok.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)13ids = tok(text, return_tensors="pt").input_ids
14out = model.generate(ids, max_new_tokens=256)15print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=False))
python
1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
34repo ="idah4/qwen3.5-4b-distilled"5tok = AutoTokenizer.from_pretrained(repo, subfolder="hf-4bit")6model = AutoModelForCausalLM.from_pretrained(7 repo, subfolder="hf-4bit",8 dtype=torch.bfloat16,9 low_cpu_mem_usage=True,10)11messages =[{"role":"user","content":"Write a Python function to merge two dictionaries."}]12text = tok.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)13ids = tok(text, return_tensors="pt").input_ids.to('cuda')14out = model.generate(ids, max_new_tokens=256)15print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=False))16
transformers >= 5.12 is recommended (native qwen3_5 support). On CPU this
runs correctly but slowly; a CUDA GPU is recommended for throughput.
For reproducible deployments, pin both the model revision and the MLX-LM
version.
Intended use
This release is intended for:
personal research;
portfolio demonstration;
local prototyping;
function-calling experiments;
agent-framework integration tests;
short coding and automation tasks.
Limitations
A 4B model is not a replacement for larger, independently evaluated coding
or agentic models.
Tool selection and arguments may be incorrect.
Generated commands and code may be insecure or destructive.
The model may call a tool unnecessarily or fail to call one when needed.
Quantization may reduce quality relative to the fused higher-precision
checkpoint.
Fine-tuning may regress capabilities of the base model.
Korean evaluation is currently limited.
Multi-turn and multi-tool behavior has not been comprehensively evaluated.
Long-context performance has not been comprehensively evaluated.
Internal evaluation results may be affected by repeated checkpoint
selection against the same validation set.
The model may reproduce errors, biases, or memorized patterns from its
training data.
Treat every generated tool call, command, argument, and code fragment as
untrusted until validated.
Security
Applications should use:
explicit tool allowlists;
strict argument schemas;
command and path validation;
sandboxed execution;
least-privilege credentials;
confirmation before destructive operations;
secret filtering;
audit logging;
resource and rate limits.
Do not expose this model directly to unrestricted shell, filesystem,
network, production infrastructure, or payment-system access.
Release and reproducibility
Each public benchmark result should be associated with:
an immutable model commit;
the exact base-model revision;
training-code revision;
dataset revisions;
MLX and MLX-LM versions;
random seed;
quantization configuration;
evaluation-code revision.
Recommended release naming:
text
1v0.1-step3900
2v0.2-stepXXXX
Existing version tags should not be overwritten.
License
The modified model weights in this repository are released under the Apache
License 2.0, consistent with the license indicated by the base-model
repository.
Modifications include:
multi-stage LoRA supervised fine-tuning;
adapter fusion;
conversion to MLX format;
export to Hugging Face bf16 safetensors for transformers/vLLM (hf-bf16/);
6-bit quantization;
chat-template and serving-parser adjustments.
This repository should retain the applicable Apache-2.0 license text,
copyright notices, attribution notices, and any applicable NOTICE material
from the base model.
Third-party training datasets remain subject to their original licenses,
terms, and attribution requirements. Inclusion in the training mixture does
not relicense those datasets under Apache-2.0.
The model and associated materials are provided on an “AS IS” basis, without
warranties or conditions of any kind, as provided by the applicable license.
Users are responsible for evaluating the model and complying with applicable
laws, dataset terms, third-party licenses, and security requirements.