Zen 6
The Zen 6 generation, available now · 27B dense · text, images and video · 1M-token context
Zen 6 is the current Zen generation from Zen LM, the open model family of Zoo Labs Foundation, a 501(c)(3) non-profit. It is chosen for two jobs:
- Agentic coding that runs on your own machine. A 27B dense model with a 1,048,576-token context, tool calling, and a bundled speculative drafter for faster decoding. The numbers below were measured on one DGX Spark.
- Marketing work. It reads images and video beside text, so the brief, the brand assets and the draft sit in one context.
Its compact build is Zen 6 Flash.
The next generation, Zen 7, is in research preview:
request access. Zen 6 answers on
api.hanzo.ai as zen6.
Specifications
Every figure below is read from this repository's config.json.
| Parameters | 27.3 billion in the language model, dense, plus a vision encoder |
| Layers | 64: 48 linear-attention (gated delta) and 16 full-attention, one full every fourth |
| Hidden size | 5,120 |
| Attention heads | 24 query, 4 key/value (6:1 grouped-query), head size 256 |
| Vocabulary | 248,320 tokens |
| Context | 262,144 tokens native; 1,048,576 with YaRN (factor 4) |
| Inputs | text, images, video |
| Weights | NVFP4 (W4A4, group size 16) on the MLP and LM head, FP8 (E4M3) on attention |
| LM head | NVFP4 holds it in about 0.7 GB where BF16 needs 2.5 GB (248,320 × 5,120 weights) |
| Speculative drafter | bundled in dflash2/, block-diffusion, 1M position table |
| Architecture | Qwen3_5ForConditionalGeneration (model_type: qwen3_5) |
Calibration
The mixed-precision weights were produced with NVIDIA ModelOpt 0.47.0.dev0:
{
"calibration_dataset": "abisee/cnn_dailymail",
"calibration_samples": 1024,
"calibration_seq_len": 512,
"algorithm": "max",
"quant_scheme": "MIXED_PRECISION",
"lm_head": { "quant_algo": "NVFP4", "group_size": 16 },
"mlp_layers": { "quant_algo": "NVFP4", "group_size": 16 },
"attention_layers": { "quant_algo": "FP8" },
"context_extension": {
"rope_type": "yarn",
"rope_theta": 10000000,
"factor": 4.0,
"partial_rotary_factor": 0.25,
"max_position_embeddings": 1048576
}
}
Speculative decoding
The drafter in dflash2/ proposes a block of 3 to 5 tokens in one forward pass,
and Zen 6 verifies the block in one parallel step, so greedy output is exactly
what Zen 6 alone would produce.
Measured on one DGX Spark
GB10 (SM121), CUDA 13.3, 128 GB unified LPDDR5X.
Prefill against context length
| Context | Cold prefill (tok/s) | Warm prefill (tok/s) | Speedup |
|---|---|---|---|
| 512 | 2,891.4 | 19,450.0 | 6.73x |
| 2,048 | 2,658.3 | 24,120.5 | 9.07x |
| 8,192 | 1,835.3 | 28,490.2 | 15.52x |
| 16,384 | 1,700.7 | 31,180.0 | 18.33x |
| 32,768 | 1,414.4 | 33,520.1 | 23.70x |
Decode with the drafter
| Task | Alone | With the drafter | Draft acceptance | Speedup |
|---|---|---|---|---|
| Code completion (Python, Rust) | 62.4 tok/s | 141.2 tok/s | 71.4% | 2.26x |
| Tool calling and JSON | 58.1 tok/s | 128.8 tok/s | 68.2% | 2.22x |
| Reasoning | 54.0 tok/s | 109.8 tok/s | 57.9% | 2.03x |
Serve it
SGLang with the drafter, from a local copy (the drafter is a folder of this repository, not a repository of its own):
hf download zenlm/zen6 --local-dir zen6
python3 -m sglang.launch_server \
--model-path ./zen6 \
--speculative-draft-model-path ./zen6/dflash2 \
--speculative-num-steps 3 \
--speculative-algorithm DFLASH \
--kv-cache-dtype fp8_e5m2 \
--context-length 1048576 \
--port 30000 \
--host 0.0.0.0
Or call it hosted, with no weights to run:
curl https://api.hanzo.ai/v1/chat/completions \
-H "Authorization: Bearer $HANZO_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "zen6", "messages": [{"role": "user", "content": "Hello"}]}'
Citation
@misc{zenlm2026zen6,
title = {Zen 6},
author = {Zen LM},
year = {2026},
publisher = {Zoo Labs Foundation}
}
License & attribution
Apache-2.0; see LICENSE. Zen 6
is built from Qwen3.8-27B by the Qwen team (Apache-2.0), with the NVFP4
quantization from RadixArk (RadixArk/Qwen3.8-27B-NVFP4) and the speculative
drafter from incoai (incoai/Qwen3.8-27B-DFlash2).