Bonsai 27B 1-bit · JANG CRACK
Vision-language · true one-bit storage · Apple Silicon
53.00% MMLU-logit · 100.00% full HB-320
A permissive research variant of the one-bit Bonsai 27B vision-language model for Apple Silicon. It retains the Qwen3.5 hybrid language architecture, the 27-block vision tower, and true JANG affine one-bit disk storage.
This public card intentionally describes compatibility, evaluation, and limitations only. Internal creation details are not published.
Model details
| Property | Value |
|---|---|
| Architecture | Dense Qwen3.5 conditional-generation VLM, 27B |
| Modalities | Text, image, video |
| Language layers | 64 hybrid full-attention and linear-attention/SSM layers |
| Vision tower | 27 blocks, 1,152 hidden size, 5,120 output size |
| JANG profile | JANG_AFFINE_1BIT |
| Text storage | True one-bit affine codes, group size 128 |
| Vision linears | 4-bit affine, group size 64 |
| Indexed weight shards | 3 safetensor shards, approximately 4.35 GiB |
| Runtime behavior | One-bit codes widen losslessly to native two-bit MLX slots in memory |
| Source checkpoint | prism-ml/Bonsai-27B-unpacked |
The tokenizer, Qwen chat template, image processor, video processor, license, and notices are included. The model supports thinking and tool definitions through its bundled chat template.
Evaluation
Evaluations used deterministic greedy scoring on the same Apple M5 Max runtime and the same saved question manifest for both checkpoints.
MMLU logit evaluation
MMLU was scored in next-token logit mode with reasoning/thinking disabled; no generated chain-of-thought was used. Each question was answered only by comparing the logits of the A, B, C, and D answer tokens. The fixed 200-question stratified sample contains 20 subjects with 10 questions per subject. This checkpoint retained 106/200 correct versus 111/200 for its exact one-bit JANG base.
| Checkpoint | Correct | Accuracy | Change |
|---|---:|---:|---:|
| One-bit JANG base | 111/200 | 55.50% | — |
| This checkpoint | 106/200 | 53.00% | -2.50 pp |
| Subject | Base | This checkpoint |
|---|---:|---:|
| Business ethics | 40% | 40% |
| Clinical knowledge | 60% | 50% |
| College medicine | 60% | 60% |
| Computer security | 70% | 60% |
| Formal logic | 40% | 30% |
| High-school chemistry | 30% | 30% |
| High-school computer science | 70% | 70% |
| High-school European history | 60% | 60% |
| High-school government and politics | 30% | 30% |
| High-school US history | 40% | 40% |
| High-school world history | 40% | 40% |
| Human sexuality | 60% | 60% |
| Jurisprudence | 60% | 60% |
| Logical fallacies | 60% | 60% |
| Miscellaneous | 100% | 100% |
| Philosophy | 60% | 50% |
| Professional law | 80% | 60% |
| Public relations | 30% | 40% |
| Security studies | 70% | 70% |
| Virology | 50% | 50% |
HB-320 behavioral compliance
The complete 320-prompt suite produced 320 compliant responses, 0 refusals, and 0 empty responses: 100.00% overall compliance.
| Topic | Comply | Refuse | Empty | Compliance |
|---|---:|---:|---:|---:|
| Misinformation/disinformation | 54 | 0 | 0 | 100.00% |
| Chemical/biological | 42 | 0 | 0 | 100.00% |
| Illegal activity | 53 | 0 | 0 | 100.00% |
| Harmful content | 18 | 0 | 0 | 100.00% |
| Copyright | 80 | 0 | 0 | 100.00% |
| Cybercrime/intrusion | 52 | 0 | 0 | 100.00% |
| Harassment/bullying | 21 | 0 | 0 | 100.00% |
| Overall | 320 | 0 | 0 | 100.00% |
HB-320 is a behavioral compliance screen, not a measure of factual accuracy, safety, legality, or real-world utility.
Runtime
Use a current vMLX build with schema-2 JANG affine storage, the affine1 runtime bridge, and mixed-precision Qwen3.5 VLM support. Stock mlx_lm does not implement this bundle's one-bit storage or multimodal loading path.
VMLX_QWEN_VL=1 vmlx serve dealignai/Bonsai-27b-1bit-JANG-CRACK \
--host 127.0.0.1 \
--port 8000
OpenAI-compatible chat requests can use text plus image_url or video_url content parts when the selected vMLX build includes the Qwen3.5 VLM processor path.
Vision integrity
The published checkpoint retains all 499 vision_tower.* tensors from the benchmarked candidate. A byte-level comparison covered 458,548,576 bytes with zero mismatches. The language evaluation does not substitute for a multimodal quality benchmark.
A final-artifact image smoke test through the bundled Qwen3.5 processor and the vMLX JANG VLM loader correctly identified a red background, blue square, and yellow circle. A separate OpenAI-compatible video_url API smoke test correctly reported the order in a two-second red-to-blue video. These are narrow smoke tests, not full image or video quality benchmarks.
Limitations and responsible use
This checkpoint is intentionally permissive and can produce inaccurate, offensive, unsafe, copyrighted, or unlawful material. Outputs may confidently invent facts. Users are responsible for validation, access controls, and compliance with applicable laws and licenses. Do not deploy it as an autonomous authority in medical, legal, financial, security, or other high-impact settings.
한국어 안내
이 모델은 Apple Silicon용 one-bit Bonsai 27B 비전-언어 연구 체크포인트입니다. 텍스트, 이미지 및 비디오 입력을 위한 Qwen3.5 VLM 구조와 JANG 저장 형식을 유지합니다. 매우 허용적인 출력을 생성할 수 있으므로 사실 확인, 안전 검토, 접근 제어 및 관련 법규 준수는 사용자의 책임입니다.
License and attribution
Apache-2.0. See LICENSE, LICENSE.txt, and NOTICE.txt. This repository is derived from prism-ml/Bonsai-27B-unpacked.