Mephisto-4B-v2.1
Abstract
Mephisto-4B-v2.1 is a 4B-parameter agentic language model constructed via multi-stage model merging using mergekit. The model integrates reasoning, coding, and agentic capabilities into a single 4B checkpoint based on the Qwen3.5-4B architecture, targeting performance exceeding Mephisto-4B-v2 on Japanese and agentic benchmarks while maintaining deployment feasibility on consumer hardware (RTX 3060 12GB).
Methodology
Merge Pipeline
The model is constructed through a three-stage merging pipeline:
Stage 1: Reasoning-Code Fusion (NuSLERP)
Normalized spherical linear interpolation (NuSLERP) on task vectors derived from a common base model (Qwen3.5-4B). This stage fuses a reasoning-distilled model (Opus-reasoning distillation) with a code-specialized model.
| Component | Source Model | Weight |
|---|
| Reasoning | Jackrong/Qwen3.5-4B-Claude-4.6-Opus-Reasoning-Distilled-v2 | 1.0 |
| Code | Jackrong/Qwopus3.5-4B-Coder | 0.7 |
Configuration:
1merge_method: nuslerp
2base_model: Qwen/Qwen3.5-4B
3models:
4 - model: Jackrong/Qwen3.5-4B-Claude-4.6-Opus-Reasoning-Distilled-v2
5 parameters: {weight: 1.0}
6 - model: Jackrong/Qwopus3.5-4B-Coder
7 parameters: {weight: 0.7}
8parameters:
9 nuslerp_flatten: true
10 nuslerp_row_wise: false
11dtype: bfloat16
Stage 2: Agent Foundation (DARE-TIES 3-way)
DARE-TIES merging with density-based pruning and sign consensus. This stage integrates three complementary agentic models into a unified foundation.
| Component | Source Model | Weight | Density |
|---|
| Reasoning | BAAI/AREX-Turbo | 1.0 | 0.5 |
| Agent/Planning | InternScience/Agents-A1-4B | 1.0 | 0.5 |
| General/Japanese | Jackrong/Qwopus3.5-4B-v3 | 0.6 | 0.5 |
Configuration:
1merge_method: dare_ties
2base_model: Qwen/Qwen3.5-4B
3models:
4 - model: BAAI/AREX-Turbo
5 parameters: {weight: 1.0, density: 0.5}
6 - model: InternScience/Agents-A1-4B
7 parameters: {weight: 1.0, density: 0.5}
8 - model: Jackrong/Qwopus3.5-4B-v3
9 parameters: {weight: 0.6, density: 0.5}
10parameters:
11 lambda: 0.8
12 density: 0.5
13dtype: bfloat16
Stage 3: Final Integration (FrankenMerge / Passthrough)
Layer-wise model selection (FrankenMerge) via Passthrough merge. This stage structurally allocates layers by functional specialization:
| Layer Range | Source | Functional Role |
|---|
| 0–15 | Qwen/Qwen3.5-4B (Instruct) | Embeddings, shallow features, Japanese language modeling |
| 16–23 | Stage 2 Output (Agent Foundation) | Planning, function calling, agentic reasoning |
| 24–31 | Stage 1 Output (Reasoning + Code) | Deep reasoning, code generation, complex logic |
Configuration:
1merge_method: passthrough
2slices:
3 - sources:
4 - model: Qwen/Qwen3.5-4B
5 layer_range: [0, 16]
6 - sources:
7 - model: merged/s2
8 layer_range: [16, 24]
9 - sources:
10 - model: merged/s1
11 layer_range: [24, 32]
Source Models
| Model | Role | Type | License |
|---|
| Jackrong/Qwen3.5-4B-Claude-4.6-Opus-Reasoning-Distilled-v2 | Reasoning distillation | Instruct | Apache-2.0 |
| Jackrong/Qwopus3.5-4B-Coder | Code / Tool Use | Instruct | Apache-2.0 |
| BAAI/AREX-Turbo | General reasoning | Base/Instruct | Apache-2.0 |
| InternScience/Agents-A1-4B | Agent / Planning / FC | Instruct | Apache-2.0 |
| Jackrong/Qwopus3.5-4B-v3 | Japanese / General | Instruct | Apache-2.0 |
| Qwen/Qwen3.5-4B | Base anchor (Instruct) | Instruct | Apache-2.0 |
All source models are Apache-2.0 licensed, enabling commercial use of the merged artifact.
Benchmarks (In Progress)
| Benchmark | Status | Category |
|---|
| ELYZA-tasks-100 | In Progress | Japanese instruction following |
| MT-Bench-JP | In Progress | Japanese conversation quality |
| JGLUE | In Progress | Japanese NLU |
| JMMLU | In Progress | Japanese knowledge |
| HumanEval | In Progress | Code generation |
| BFCL (Function Calling) | In Progress | Function calling |
| JSB (Safety) | In Progress | Safety alignment |
Benchmark evaluation is currently in progress. Target thresholds are based on Mephisto-4B-v2 baseline and merge methodology expectations.
Quantization
| Format | Size | Command |
|---|
| Q8_0 | 4.5 GB | llama.cpp convert_hf_to_gguf.py --outtype q8_0 --no-mtp |
| Q4_K_M | ~2.6 GB | llama.cpp convert_hf_to_gguf.py --outtype q4_k_m --no-mtp |
Note: The --no-mtp flag is required to exclude Multi-Token Prediction (MTP) modules present in Qwen3.5-series models, which are incompatible with llama.cpp inference.
Usage
HF Transformers
1from transformers import AutoModelForCausalLM, AutoTokenizer
2model = AutoModelForCausalLM.from_pretrained("Mephisto-4B-v2.1", torch_dtype="bfloat16", device_map="auto")
3tokenizer = AutoTokenizer.from_pretrained("Mephisto-4B-v2.1")
llama.cpp (Recommended)
1./llama-cli -m Mephisto-4B-v2.1_q8_0.gguf -ngl 99 -c 4096 -p "ユーザー: こんにちは
2アシスタント: "
Hardware Requirements
| Component | Specification |
|---|
| GPU | RTX 3060 12GB+ (for Q8_0 with -ngl 99) |
| RAM | 16 GB+ system RAM |
| Disk | ~12 GB (HF + GGUF) |
License
This model inherits licenses from its source models. All source models are Apache-2.0 licensed, permitting commercial use. Users should verify individual source model licenses for their specific use cases.
Citation
If you use this model in your research, please cite:
1@misc{mephisto-4b-v2.1,
2 title={Mephisto-4B-v2.1: Multi-Stage Merged Agentic Model},
3 author={CloudGoat},
4 year={2024},
5 note={Constructed via mergekit multi-stage merging: NuSLERP + DARE-TIES + FrankenMerge}
6}
Model Card Version: 1.0
Last Updated: 2024-08-05
Merge Tool: mergekit (NuSLERP + DARE-TIES + Passthrough)
Base Architecture: Qwen3.5-4B (32 layers, 2560 hidden dim, 16/4 attention heads)