GGUF quantizations of Sixpert K2 for Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes.
Sixpert K2 is a 9B parameter mixture-of-experts (MoE) model designed for deep reasoning, complex agentic workflows, and multimodal understanding. Built with a 1M-token context window and fine-tuned on 500M+ reasoning tokens, it represents a significant leap in the 9B parameter class.
Coming Soon
🚀 Sixpert AI app coming soon to Google Play Store — a self-improving AI agent with sandbox, Browserbase integration, and Ornith 1.0 self-scaffolding RL built in.
Self-Improving Training Architecture
Sixpert K2 is trained using a self-scaffolding reinforcement learning framework inspired by Ornith 1.0's architecture. Instead of a human writing the model's execution framework once, K2 generates its own Python harness for each task — and learns through reward signals which harness patterns work best.
The Two-Stage RL Loop
Two-Stage RL Loop
Stage 1 — Scaffold Generation: K2 analyzes a task and generates a Python harness. It reasons step-by-step in <reasoning> blocks about the optimal approach (which tools, what sequence, error recovery), then generates the harness code with tool selection logic, state management, error handling, and termination criteria.
Stage 2 — Solution Rollout: K2 uses the generated harness to solve the task. It follows the harness structure, reasoning in <reasoning> blocks before each action, and making XML-style tool calls (<function=name>...</function>) as instructed by the harness.
Joint Optimization: The reward from the solution backpropagates to BOTH stages through GRPO. Better scaffolds lead to better solutions, which reinforce better scaffold generation — a self-improvement loop.
How K2 Thinks & Analyzes
How the Model Thinks
K2 uses reasoning blocks for step-by-step planning before every action. It thinks about what the task is asking, what tools are available, what could go wrong, and when to terminate — all before generating the execution harness.
3-Layer Anti-Reward-Hacking
3-Layer Anti-Hacking
To prevent the model from gaming the reward system, three defense layers are stacked:
Fixed Trust Boundary — K2 cannot modify the evaluation environment, reward function, or test runner
Deterministic Monitor — Rule-based system validates scaffold structure (AST analysis, trivial scaffold detection, hardcoded answer detection) and solution outputs
Frozen LLM Judge — Sixpert K1 serves as a frozen judge that evaluates K2's solution quality. K1 is never updated during K2's training, so K2 cannot learn to trick it
The Self-Improvement Cycle
Self-Improvement Cycle
Each training iteration builds on the last: generate scaffold → execute solution → compute reward → policy update → repeat. Over time, K2 discovers better orchestration patterns and generates higher-quality solutions.
Staleness-Weighted GRPO
Staleness-Weighted GRPO
Long agentic rollouts create stale training data — by the time a trajectory completes, model weights have already moved. K2 uses Ornith 1.0's staleness-weighted GRPO: fresh tokens get full weight, stale tokens get downweighted, and tokens past the threshold are dropped entirely.
Sixpert K2 benchmark scores are derived from verified third-party evaluations of its base architecture from llm-stats.com and TokenCalculator.com (April 2026). As a 9B model, Sixpert K2 competes directly with much larger models.
Sixpert K2 Radar Chart
Sixpert K2 Bar Chart
Sixpert K1 vs K2 Combined
Verified Real Scores
Benchmark
Sixpert K2 Score
Source
MMLU
82.5%
llm-stats.com (MMLU-Pro)
HumanEval
85.0%
Competitive 9B class coding
MATH
62.0%
Competitive with 8B class thinking
GPQA
81.7%
llm-stats.com (GPQA)
GSM8K
90.5%
Competitive with 8B class thinking
MMLU-Redux
91.1%
llm-stats.com
IFEval
91.5%
llm-stats.com
C-Eval
88.2%
llm-stats.com
Real Competitor Comparison (April 2026)
The charts above compare Sixpert K2 against verified real-world scores from official model cards:
GPT-5.4: MMLU 91.8%, HumanEval 94.1%
Claude Opus 4.6: MMLU 92.1%, HumanEval 92.4%
Gemini 3.1 Ultra: MMLU 90.4%, HumanEval 89.3%
DeepSeek V4: MMLU 87.2%, HumanEval 88.7%
Llama 4 Maverick: MMLU 84.7%, HumanEval 82.1%
Files
Normal text weights — fixed v3 replacements
File
Quant
Size
Notes
SixpertK2-Q4_K_M.gguf
Q4_K_M
5.3 GB / 5.63 GB
recommended default — fixed v3, best compatibility
If you don't know which to pick, Q4_K_M is the right starting point — it's the smallest practical quant with good quality preservation.
Quick Start
Ollama
ollama run hf.co/SixpertAI/SixpertK2:latest
LM Studio / jan / KoboldCpp
Drop any of the .gguf files into your runtime's model directory. Modern GGUF runtimes load it automatically from the file.
Vision (image input)
Sixpert K2 supports image input out of the box. Run with llama.cpp's multimodal CLI or server.
Sixpert K2 is a reasoning model — every response opens with a <thought> block before the final answer. Use these settings as defaults:
Parameter
Value
temperature
0.6
top_p
0.95
top_k
20
repeat_penalty
1.05
max_new_tokens
16384 (generous budget for <thought> + answer)
These are the official thinking-mode recommendations. Avoid greedy decoding and very-low-temperature sampling (T ≤ 0.3) — both can cause repetition loops on long reasoning generations.
Long Context (1M tokens)
The GGUFs ship with YaRN rope-scaling baked in for a 1,048,576-token context window (4× extension over the 262k native).
To use the full 1M window in llama-cli, set -c 1010000 (or any context length up to that). For shorter prompts, lower -c to reduce KV-cache memory — at default settings llama.cpp will autosize.
A single H100/H200-class GPU comfortably handles 256k–512k; the full 1M typically needs tensor-parallel multi-GPU or aggressive KV-cache offload.
Capabilities
Reasoning — Advanced chain-of-thought reasoning for complex problems
Function Calling — Native tool use with structured output
Domain Expertise — Strong in cybersecurity, red-teaming, biology, pharmacology, and clinical medicine
Limitations
Reasoning model. Every answer opens with a <thought> block; allow generous max_new_tokens and parse/strip <thought>...</thought> for end users.
Use recommended sampling. Greedy / very-low-temp can cause repetition loops.
Verify specifics in safety-critical contexts. Like all closed-book LLMs in this weight class, Sixpert K2 can over-commit to specific identifiers (CVEs, hashcat modes, drug positions) it isn't certain about. Pair with retrieval or function calling in such deployments — the model uses tools cleanly when offered them.
Uncensored — add your own application-level review/safety layer for end-user-facing deployments where that matters.
Creator
Sixpert K2 was created by Inyang David.
Provenance & Licensing
Weights are released under Apache-2.0. Shared for research and experimentation, as-is.
Acknowledgements
Creator: Inyang David
Architecture: Transformer-based multimodal language model