zenlm/zen6

🤗 Hugging Face sourceimage-text-to-textapache-2.018.2B params21 GBsafetensors✓ 5 checksumsupdated today
Needs seeder →

Zen 6

The Zen 6 generation, available now · 27B dense · text, images and video · 1M-token context

Zen 6 is the current Zen generation from Zen LM, the open model family of Zoo Labs Foundation, a 501(c)(3) non-profit. It is chosen for two jobs:

  • Agentic coding that runs on your own machine. A 27B dense model with a 1,048,576-token context, tool calling, and a bundled speculative drafter for faster decoding. The numbers below were measured on one DGX Spark.
  • Marketing work. It reads images and video beside text, so the brief, the brand assets and the draft sit in one context.

Its compact build is Zen 6 Flash. The next generation, Zen 7, is in research preview: request access. Zen 6 answers on api.hanzo.ai as zen6.

Specifications

Every figure below is read from this repository's config.json.

Parameters 27.3 billion in the language model, dense, plus a vision encoder
Layers 64: 48 linear-attention (gated delta) and 16 full-attention, one full every fourth
Hidden size 5,120
Attention heads 24 query, 4 key/value (6:1 grouped-query), head size 256
Vocabulary 248,320 tokens
Context 262,144 tokens native; 1,048,576 with YaRN (factor 4)
Inputs text, images, video
Weights NVFP4 (W4A4, group size 16) on the MLP and LM head, FP8 (E4M3) on attention
LM head NVFP4 holds it in about 0.7 GB where BF16 needs 2.5 GB (248,320 × 5,120 weights)
Speculative drafter bundled in dflash2/, block-diffusion, 1M position table
Architecture Qwen3_5ForConditionalGeneration (model_type: qwen3_5)

Calibration

The mixed-precision weights were produced with NVIDIA ModelOpt 0.47.0.dev0:

{
  "calibration_dataset": "abisee/cnn_dailymail",
  "calibration_samples": 1024,
  "calibration_seq_len": 512,
  "algorithm": "max",
  "quant_scheme": "MIXED_PRECISION",
  "lm_head": { "quant_algo": "NVFP4", "group_size": 16 },
  "mlp_layers": { "quant_algo": "NVFP4", "group_size": 16 },
  "attention_layers": { "quant_algo": "FP8" },
  "context_extension": {
    "rope_type": "yarn",
    "rope_theta": 10000000,
    "factor": 4.0,
    "partial_rotary_factor": 0.25,
    "max_position_embeddings": 1048576
  }
}

Speculative decoding

The drafter in dflash2/ proposes a block of 3 to 5 tokens in one forward pass, and Zen 6 verifies the block in one parallel step, so greedy output is exactly what Zen 6 alone would produce.

Measured on one DGX Spark

GB10 (SM121), CUDA 13.3, 128 GB unified LPDDR5X.

Prefill against context length

Context Cold prefill (tok/s) Warm prefill (tok/s) Speedup
512 2,891.4 19,450.0 6.73x
2,048 2,658.3 24,120.5 9.07x
8,192 1,835.3 28,490.2 15.52x
16,384 1,700.7 31,180.0 18.33x
32,768 1,414.4 33,520.1 23.70x

Decode with the drafter

Task Alone With the drafter Draft acceptance Speedup
Code completion (Python, Rust) 62.4 tok/s 141.2 tok/s 71.4% 2.26x
Tool calling and JSON 58.1 tok/s 128.8 tok/s 68.2% 2.22x
Reasoning 54.0 tok/s 109.8 tok/s 57.9% 2.03x

Serve it

SGLang with the drafter, from a local copy (the drafter is a folder of this repository, not a repository of its own):

hf download zenlm/zen6 --local-dir zen6

python3 -m sglang.launch_server \
  --model-path ./zen6 \
  --speculative-draft-model-path ./zen6/dflash2 \
  --speculative-num-steps 3 \
  --speculative-algorithm DFLASH \
  --kv-cache-dtype fp8_e5m2 \
  --context-length 1048576 \
  --port 30000 \
  --host 0.0.0.0

Or call it hosted, with no weights to run:

curl https://api.hanzo.ai/v1/chat/completions \
  -H "Authorization: Bearer $HANZO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "zen6", "messages": [{"role": "user", "content": "Hello"}]}'

Citation

@misc{zenlm2026zen6,
  title  = {Zen 6},
  author = {Zen LM},
  year   = {2026},
  publisher = {Zoo Labs Foundation}
}

License & attribution

Apache-2.0; see LICENSE. Zen 6 is built from Qwen3.8-27B by the Qwen team (Apache-2.0), with the NVFP4 quantization from RadixArk (RadixArk/Qwen3.8-27B-NVFP4) and the speculative drafter from incoai (incoai/Qwen3.8-27B-DFlash2).