sakamakismile/Qwen3.6-27B-Fable-Fusion-MTP-NVFP4

🤗 Hugging Face 来源image-text-to-textapache-2.016.7B 参数20 GBsafetensors✓ 3 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo sakamakismile/Qwen3.6-27B-Fable-Fusion-MTP-NVFP4 ./model-folder
需要做种者 →

Qwen3.6-27B-Fable-Fusion-MTP-NVFP4

Publication status: P2c quality gates passed; approved for Hugging Face publication. The required runtime contract is to leave --kv-cache-dtype unset so vLLM uses auto / the model's BF16 dtype. Do not use FP8 KV cache with the validated vLLM release.

This is an SM120-oriented NVFP4 conversion of nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-F32, an Apache-2.0 NuSLERP merge of Qwen3.6-27B-Architect-Polaris2-Fable-B and Qwen3.6-27B-Architect-Polaris-Fable-F451. The language-model weights were quantized directly from the F32 safetensors original. No GGUF, MLX, Q8, or pre-existing NVFP4 checkpoint was used as a quantization source.

Conversion

  • Recipe: llmcompressor W4A4 NVFP4, targets="Linear".
  • Calibration: 512 deterministic mixed samples, 4,096-token maximum, seed 42.
  • Composition: 64 each of Japanese dialogue, English dialogue, Python, TypeScript, shell, JSON tool calls, and long reasoning; 32 each of creative writing and vision-shaped dialogue.
  • All 512 rendered calibration samples contain an empty Thinking block; none contains non-empty reasoning inside that block. This is a known coverage gap. P2c did not rebake because the cheaper controlled KV-cache test isolated and fixed the regression, and every static, capability, Thinking, acceptance, and performance gate then passed on the unchanged checkpoint.
  • Quantized modules: 496 language-model Linear modules.
  • Runtime format: compressed-tensors, nvfp4-pack-quantized.
  • Payload size: 20,558,935,392 bytes in two safetensors shards.
  • Physical model files: about 20 GB.

The following ranges remain BF16:

  • Vision tower: 333 tensors.
  • MTP head: 15 tensors, 849,398,784 payload bytes.
  • lm_head: 1 tensor.
  • conv1d / Gated DeltaNet special layers: 48 tensors.
  • Norms and other unsupported or non-target tensors remain unquantized as selected by the canonical recipe.

All 15 MTP module names are explicitly present in quantization_config.ignore. Omitting these entries is known to produce 0% MTP acceptance on this architecture.

MTP origin decision

The input merge contains 15 mtp.* tensors: mtp.fc.weight is BF16 and the other 14 are F32. They were cast directly to BF16 and compared against a self-extracted BF16 MTP head from Qwen/Qwen3.6-27B.

  • mtp.fc.weight was exactly identical.
  • The other 14 tensors differed after BF16 cast.
  • Cosine similarities were approximately 0.999996–0.999998.
  • Maximum absolute difference was at most 0.0078125.

Matched TP=2, MTP k=2 A/B testing selected B, the official origin:

MTP origin Draft acceptance First-position acceptance Output tok/s
A — input-derived 67.34% 78.72% 42.17
B — official 70.53% 79.85% 43.23

B improved full-draft acceptance by 3.19 percentage points and tok/s by 2.51% on the fixed eight-prompt single-stream workload. The final local directory therefore contains the official B MTP payload. Static verification confirms 15/15 BF16 tensors, exact expected payload hashes, index correctness, explicit ignore registration, and an unchanged language-model shard SHA-256 versus BASE.

vLLM startup

The measured runtime was vllm/vllm-openai:v0.26.0, TP=2. On this host, NCCL_P2P_DISABLE=1 and --disable-custom-all-reduce are required. Select two compute GPUs and do not include the display GPU. Leave --kv-cache-dtype unset: vLLM resolves it to auto with BF16 for this model. FP8 KV cache is not a supported launch configuration for this release.

MODEL=/path/to/Qwen3.6-27B-Fable-Fusion-MTP-NVFP4
IMAGE=vllm/vllm-openai:v0.26.0

serve() {
  docker run --rm \
    --gpus '"device=0,1"' \
    --ipc=host \
    --shm-size=8g \
    -e NCCL_P2P_DISABLE=1 \
    -p 8000:8000 \
    -v "$MODEL:$MODEL:ro" \
    --entrypoint python3 \
    "$IMAGE" \
    -m vllm.entrypoints.openai.api_server \
    --model "$MODEL" \
    --served-model-name fable-fusion \
    --trust-remote-code \
    --tensor-parallel-size 2 \
    --disable-custom-all-reduce \
    --max-model-len 32768 \
    --limit-mm-per-prompt '{"image":1,"video":0}' \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_xml \
    --host 0.0.0.0 \
    --port 8000 \
    "$@"
}

Normal decoding / BASE mode:

serve

MTP k=2:

serve --speculative-config \
  '{"method":"mtp","num_speculative_tokens":2}'

Suffix example:

serve --speculative-config \
  '{"method":"suffix","num_speculative_tokens":8,"suffix_decoding_max_tree_depth":24,"suffix_decoding_max_cached_requests":10000,"suffix_decoding_max_spec_factor":1.0,"suffix_decoding_min_token_prob":0.1}'

N-gram example:

serve --speculative-config \
  '{"method":"ngram","num_speculative_tokens":4,"prompt_lookup_min":2,"prompt_lookup_max":5}'

Suffix and N-gram commands are configuration examples, not P2-validated recommendations.

Thinking termination and TP=2 measurements

P2c compared the identical checkpoint and vLLM path with only the KV-cache contract changed. With FP8 KV cache, the mandatory greedy 17×19 prompt used all 16,384 completion tokens without </think> or a final answer. With the flag omitted (auto, model BF16), it naturally terminated after 328 Thinking tokens and answered 323. All five seeded greedy probes naturally terminated and answered correctly:

Probe Thinking tokens Finish
17×19 328 stop
23×17 853 stop
144÷12 171 stop
1000−637 218 stop
248+175 797 stop

The specified sampling control also passed at 223 Thinking tokens. The F32 source control had previously terminated the original prompt at 303 Thinking tokens / 309 generated tokens. This controlled result identifies FP8 KV cache, not token budget or the NVFP4 weight bake, as the cause of the P2b regression.

The fixed single-stream workload generated 1,280 completion tokens over eight requests after warm-up with BF16/auto KV cache.

Mode Output tok/s Peak VRAM/GPU Relative to BASE
Normal decode 26.73 14,453 MiB 1.00x
MTP k=2 44.10 14,159 MiB 1.65x

P2c k=2 speculative metrics:

  • Draft acceptance: 68.68%.
  • First-position acceptance: 79.00%.
  • Second-position acceptance: 58.36%.
  • Mean acceptance length including bonus: 2.37.

Absolute capability checks passed for both normal and MTP runs: GSM8K-style 8/8, HumanEval-style 4/4, JSON tool call, Japanese QA 5/5, red/blue image understanding, and 8,340-token long-context retrieval.

Known limitations

  • FP8 KV cache is incompatible with the validated release contract. Under vLLM 0.26.0 it reproducibly caused greedy Thinking non-termination on the mandatory prompt. Omit the flag and verify the engine logs kv_cache_dtype=auto.
  • The calibration corpus has no non-empty text inside its 512 Thinking blocks. P2c did not rebuild it because the controlled cache-dtype change restored every required gate on the existing weight artifact.
  • P2c BASE-vs-MTP comparison matched 3/6 greedy token sequences exactly and 6/6 semantically, with zero major incompatibilities.
  • MTP k=1 and k=3 were not rerun in P2c; k=2 is the validated recommendation.
  • vLLM warns that fused parallel NVFP4 layers have different weight global scales and may reduce accuracy. The warning remains with BF16 KV cache, but all P1, P2, P2b, and static gates passed; the canonical recipe was not altered.
  • Measurements used two RTX PRO 2000 Blackwell 16 GB GPUs at a 70 W power cap, concurrency 1, a 16,384-token P2 quality server, and a 32,768-token P2b server. Results are not transferable without remeasurement.
  • Suffix and N-gram were not benchmarked in P2.
  • DFlash・詳細ベンチは続報.

Reproduction and evidence

The project contains:

  • src/mtp/extract_official_mtp.py
  • src/mtp/graft_mtp.py
  • src/mtp/verify_mtp.py
  • src/benchmark/run_p2_quality_workload.py
  • src/benchmark/run_p2b_thinking_probe.py
  • src/benchmark/run_p2b_f32_cpu_probe.py
  • src/benchmark/select_mtp_candidate.py
  • src/benchmark/aggregate_p2_quality.py
  • src/benchmark/aggregate_p2c_quality.py
  • src/serve/run_p2_vllm_suite.sh
  • src/serve/run_p2c_thinking_probe.sh
  • reports/mtp_manifest.json
  • reports/quality_gate.json
  • reports/p2c_quality_gate.json
  • reports/P2_REPORT.md
  • reports/P2B_REPORT.md
  • reports/P2C_REPORT.md

License

Apache License 2.0, inherited from the source model card. Users must also follow the licenses and terms of the listed upstream base models.