Tested hardware: one Intel Arc Pro B70, 32 GB VRAM, on Linux. No B50, Windows, CPU, or offload performance claim.
Latest uncached 24K prefill:4,841 → 5,653 tokens/s (+16.8%) in a matched Heretic-reference comparison; the clean deployment build confirmed 5,667 tokens/s. Prefix reuse was OFF, kernels were warmed, and PP means input tokens divided by time to first token. This is not a cache-hit benchmark or a 6K+ result at 24K.
Shorter prompts / generation: the matched test retained about 7.4K PP at 6,622 input tokens and 112 output tokens/s on short prompts. These are reference-model figures, not guaranteed speeds for this exact fine-tune or every RP request.
Memory: about 26.8 GiB sampled peak card usage at 24K with an 8 GiB KV pool. Download size is not total VRAM. A 12 GB minimum has not been established.
Custom fork required for this gain:Wondernuttz/OpenVINO, wide-query commit, plus matching OpenVINO GenAI/tokenizers. This combines custom tuning with credited Intel/OpenVINO kernels and fixes; it is not an unmodified stock wheel.
Inference: use the complete OpenVINO export with OpenVINO GenAI, not direct Transformers loading. The tested 26B/B70 profile uses PA, DQ128, U4 KV, cache8 GiB, batch16384 and one sequence. Set GEMMA_MIXED_512_TILE=wideq before Python starts, alongside MOE_USE_GROUPED_GEMM_PREFILL=1 and MOE_GROUPED_BINARY_LOOKUP=1. RP prefix caching may stay enabled. See full instructions and rollback.
Quality / scope: all 17 matched Heretic output strings were identical; Chimera passed separate 24K recall and prefix-update checks. The clean kernel passed 24 repeated U4 reference tests with bit-identical repeats. The old baseline showed U4 variability, and an extra OFF/ON comparison gate did not pass; that caveat remains documented. This is bounded validation, not a universal intelligence guarantee.
Context and installation: the live Heretic bots received the new kernel while keeping their existing 16K context and settings. Valid exports need no recompression. Rebuild custom runtime wheels/container to get the new kernel—git pull alone does not replace OpenVINO inside an image. No new prebuilt image is supplied. 131K RoPE LUT coverage is not validated 131K usable context; reserve output space.
Chimera-X: passed separate 24K recall, RP inspection and prefix-fact-update tests with the wide-query kernel. The headline matched speed belongs to Heretic, not Chimera. This bot rollout did not change the buddy card or its selected model.
Earlier binary-lookup results (including 7,435 PP at 6,622 and 6,866 PP at 15,872) remain historical measurements with their original profiles. The wider-query result above is an additional cached/chunked-prefill improvement, not a replacement claim for every model or context size. Original model/merge credits, license and detailed historical notes below remain unchanged.
Detailed September runtime notes, configuration, and validation history
September 6 runtime update: faster long-prompt processing
Clean Release follow-up and rollout status
This checkpoint completed the clean Release 24K factual/boundary/prefix-cache gates and is included in the local runtime rollout. Manual RP review found broadly coherent prose, not perfect instruction following; see the per-model caveats.
The clean Release report records a matched Heretic 24K result of 2,711 → 4,771 PP tok/sec (+76%), the per-checkpoint decode/cache measurements, exact runtime hashes, and limitations. The kernel journal also recorded engine resets when the temporary canary server was terminated after its successful replies; shutdown/model-switch stability is not certified. The old runtime was retained for rollback. These updates change runtime documentation, not model weights or repository visibility.
The numbers below are reference-stack measurements on Gemma 4 26B A4B Heretic, not a benchmark of this exact fine-tune. This checkpoint belongs to the related Gemma 26B A4B MoE family; per-checkpoint speed, memory, and quality need their own validation.
The Wondernuttz OpenVINO fork update and full report adds an opt-in binary search for grouped-MoE token-row lookup and incorporates eight credited Intel upstream backports. No model reconversion or weight download is required to get this runtime optimization. It requires the patched, ABI-matched OpenVINO runtime; an environment variable alone cannot add it to a stock wheel.
Heretic reference input
Previous lookup
New lookup
Matched PP gain
6,622 tokens
6,028 PP tok/sec
7,435 PP tok/sec
23%
15,872 tokens
2,921 PP tok/sec
6,866 PP tok/sec
135% / 2.35x
24,576 tokens
2,735 PP tok/sec
4,767 PP tok/sec
74%
One Intel Arc Pro B70 (32 GB), DQ128, U4 KV, 8 GiB cache allocation, batch16384, one sequence, prefix reuse off. PP is input tokens divided by TTFT. At 24K, TTFT was ~5.15 seconds, long-context decode ~93 tok/sec, sampled VRAM ~26.5 GiB. Decode was essentially unchanged. The 24K matched pair produced identical output strings and passed 11 automated checks; this is bounded coherence evidence, not a universal intelligence or production-stability guarantee.
With prefix caching on, an exact 24K repeat reached first token in 173 ms. That is cache-hit latency, not raw prefill throughput. An updated-fact probe retrieved the corrected password but still answered earlier embedded questions, so strict output-only formatting was not perfect. Separate patched-only 32K probes also passed; do not automatically raise this fine-tune's serving cap to match a reference experiment.
To opt in, set these before pipeline creation, using the matching fork runtime:
The report includes the PA scheduler configuration, exact revisions, binary hash, matched methodology, limitations, and rollback switch. The published measurements use a Release/O3 binary with profiling capability compiled in but measurement counters disabled; clean-release deployment validation is separate. Native vision/audio are not validated by these text tests. Preserve your model's chat template, reasoning framing, and sampler.
Correction to older tuning attribution: MOE_MICRO_GEMM_N_HINT is bypassed by the active grouped path and must not be credited as a grouped tile optimization. Older benchmark/setup sections below document their historical builds, not updated limits or identical per-tune guarantees. The new lookup does not apply to dense Gemma 31B.
Original base-model, fine-tune, merge credits and licenses below are unchanged. Wondernuttz's contribution is the OpenVINO conversion/runtime work and testing.
OpenVINO INT4 conversion of Vortex5/Chimera-X-26B-A4B, optimized for fast local inference on Intel Arc GPUs.
This repository is a quantized deployment artifact. The original fine-tune and merge work belongs to Vortex5 and the creators of its component models; this conversion does not claim authorship of that work.
Model and credits
Chimera-X is a Gemma 4 26B-A4B mixture-of-experts model intended for roleplay, creative writing, storytelling, and brainstorming. The source model combines work from:
See the source model card for its full description, merge history, intended use, and original credits.
OpenVINO conversion
Weight format: asymmetric INT4
Compression: AWQ, group size 64, ratio 1.0
MoE routing: router layers excluded from INT4 weight compression
Graph: multimodal Gemma 4 OpenVINO IR (VLMPipeline)
RoPE lookup-table optimization: 131,072 positions, clamped at index 131,071
Source model context declaration: 262,144 tokens
Serving window validated on this build: 16K tokens
32K serving is experimental and depends on KV-cache budget and runtime configuration
The 131,072-position RoPE table is the hard ceiling of this particular graph, even though the source configuration declares a larger window.
Measured Intel Arc performance
Single-request measurements on one Intel Arc Pro B70 using the custom OpenVINO 2026.4 PA/XMX runtime used by Arcanaeum:
Test
Result
Prompt processing
about 4,967 tokens/s
Short decode
about 102.8 tokens/s
Decode with about 6.7K tokens of context
about 87.7 tokens/s
First token with about 6.7K tokens of context
about 1.45 seconds
Results are workload- and runtime-dependent. These numbers are measurements of this deployment, not guarantees for every OpenVINO build or Intel GPU.
Inference with OpenVINO GenAI
Install a recent OpenVINO GenAI build and Hugging Face Hub client:
pip install -U openvino-genai huggingface-hub
Download the full repository and point VLMPipeline at the local snapshot:
python
1import openvino_genai as ov_genai
2from huggingface_hub import snapshot_download
34model_dir = snapshot_download("Wondernutts/Chimera-X-26B-A4B-int4-ov")56pipe = ov_genai.VLMPipeline(7 model_dir,8"GPU",9 DYNAMIC_QUANTIZATION_GROUP_SIZE=128,10)1112config = ov_genai.GenerationConfig()13config.max_new_tokens =51214config.do_sample =True15config.temperature =0.916config.top_p =0.9517config.repetition_penalty =1.218config.apply_chat_template =False1920prompt =(21"<bos>"22"<|turn>system\nYou are a vivid, consistent roleplay partner.<turn|>\n"23"<|turn>user\nWrite a short scene in a moonlit inn.<turn|>\n"24"<|turn>model\n"25"<|channel>thought\n<channel|>"26)2728result = pipe.generate(prompt, generation_config=config)29print(result)
The final empty thought channel disables visible reasoning for direct roleplay responses. To use the model's reasoning mode, remove that final pre-closed channel and budget additional generation tokens. Applications should hide internal reasoning from end users.
The model artifact includes vision embeddings. Native image input requires the multimodal VLMPipeline API and an OpenVINO GenAI build compatible with this Gemma 4 export. Text inference is the path benchmarked above.
Notes
This is not a Transformers checkpoint; use OpenVINO/OpenVINO GenAI rather than AutoModelForCausalLM.
The model is uncensored/creative by design. Deployers remain responsible for prompts, outputs, applicable law, and platform policy.
License: Apache-2.0, inherited from the source model. Review the upstream card and component licenses before redistribution or commercial deployment.