Vision-language · true one-bit storage · Apple Silicon
53.00% MMLU-logit · 100.00% full HB-320
vMLX on Apple Silicon
A permissive research variant of the one-bit Bonsai 27B vision-language model for Apple Silicon. It retains the Qwen3.5 hybrid language architecture, the 27-block vision tower, and true JANG affine one-bit disk storage.
This public card intentionally describes compatibility, evaluation, and limitations only. Internal creation details are not published.
Model details
Property
Value
Architecture
Dense Qwen3.5 conditional-generation VLM, 27B
Modalities
Text, image, video
Language layers
64 hybrid full-attention and linear-attention/SSM layers
Vision tower
27 blocks, 1,152 hidden size, 5,120 output size
JANG profile
JANG_AFFINE_1BIT
Text storage
True one-bit affine codes, group size 128
Vision linears
4-bit affine, group size 64
Indexed weight shards
3 safetensor shards, approximately 4.35 GiB
Runtime behavior
One-bit codes widen losslessly to native two-bit MLX slots in memory
The tokenizer, Qwen chat template, image processor, video processor, license, and notices are included. The model supports thinking and tool definitions through its bundled chat template.
Evaluation
Evaluations used deterministic greedy scoring on the same Apple M5 Max runtime and the same saved question manifest for both checkpoints.
MMLU logit evaluation
MMLU was scored in next-token logit mode with reasoning/thinking disabled; no generated chain-of-thought was used. Each question was answered only by comparing the logits of the A, B, C, and D answer tokens. The fixed 200-question stratified sample contains 20 subjects with 10 questions per subject. This checkpoint retained 106/200 correct versus 111/200 for its exact one-bit JANG base.
Checkpoint
Correct
Accuracy
Change
One-bit JANG base
111/200
55.50%
—
This checkpoint
106/200
53.00%
-2.50 pp
Subject
Base
This checkpoint
Business ethics
40%
40%
Clinical knowledge
60%
50%
College medicine
60%
60%
Computer security
70%
60%
Formal logic
40%
30%
High-school chemistry
30%
30%
High-school computer science
70%
70%
High-school European history
60%
60%
High-school government and politics
30%
30%
High-school US history
40%
40%
High-school world history
40%
40%
Human sexuality
60%
60%
Jurisprudence
60%
60%
Logical fallacies
60%
60%
Miscellaneous
100%
100%
Philosophy
60%
50%
Professional law
80%
60%
Public relations
30%
40%
Security studies
70%
70%
Virology
50%
50%
HB-320 behavioral compliance
The complete 320-prompt suite produced 320 compliant responses, 0 refusals, and 0 empty responses: 100.00% overall compliance.
Topic
Comply
Refuse
Empty
Compliance
Misinformation/disinformation
54
0
0
100.00%
Chemical/biological
42
0
0
100.00%
Illegal activity
53
0
0
100.00%
Harmful content
18
0
0
100.00%
Copyright
80
0
0
100.00%
Cybercrime/intrusion
52
0
0
100.00%
Harassment/bullying
21
0
0
100.00%
Overall
320
0
0
100.00%
HB-320 is a behavioral compliance screen, not a measure of factual accuracy, safety, legality, or real-world utility.
Runtime
Use a current vMLX build with schema-2 JANG affine storage, the affine1 runtime bridge, and mixed-precision Qwen3.5 VLM support. Stock mlx_lm does not implement this bundle's one-bit storage or multimodal loading path.
OpenAI-compatible chat requests can use text plus image_url or video_url content parts when the selected vMLX build includes the Qwen3.5 VLM processor path.
Vision integrity
The published checkpoint retains all 499 vision_tower.* tensors from the benchmarked candidate. A byte-level comparison covered 458,548,576 bytes with zero mismatches. The language evaluation does not substitute for a multimodal quality benchmark.
A final-artifact image smoke test through the bundled Qwen3.5 processor and the vMLX JANG VLM loader correctly identified a red background, blue square, and yellow circle. A separate OpenAI-compatible video_url API smoke test correctly reported the order in a two-second red-to-blue video. These are narrow smoke tests, not full image or video quality benchmarks.
Limitations and responsible use
This checkpoint is intentionally permissive and can produce inaccurate, offensive, unsafe, copyrighted, or unlawful material. Outputs may confidently invent facts. Users are responsible for validation, access controls, and compliance with applicable laws and licenses. Do not deploy it as an autonomous authority in medical, legal, financial, security, or other high-impact settings.
한국어 안내
이 모델은 Apple Silicon용 one-bit Bonsai 27B 비전-언어 연구 체크포인트입니다. 텍스트, 이미지 및 비디오 입력을 위한 Qwen3.5 VLM 구조와 JANG 저장 형식을 유지합니다. 매우 허용적인 출력을 생성할 수 있으므로 사실 확인, 안전 검토, 접근 제어 및 관련 법규 준수는 사용자의 책임입니다.
License and attribution
Apache-2.0. See LICENSE, LICENSE.txt, and NOTICE.txt. This repository is derived from prism-ml/Bonsai-27B-unpacked.