Qwen3.8-27B-heretic-ara-DFlash2
SpecForge-tuned DFlash 2 drafter for
heretic-org/Qwen3.8-27B-heretic-ara.
Target-specific — reads the target's hidden states, so use this only with
heretic-ara. For base Qwen/Qwen3.8-27B, use
z-lab/Qwen3.8-27B-DFlash2.
Trained on
- Target:
heretic-org/Qwen3.8-27B-heretic-ara - Method: SpecForge, warm-started from
z-lab/Qwen3.8-27B-DFlash2 - Data: synthetic multi-turn corpus matching production traffic
(
reasoning_effort=low,temp 1.0 / top_p 0.95 / top_k 20, tool-call episodes) - Steps: 400 (selected from an 800-step run; step 400 wins on real vLLM bench)
vs stock z-lab/Qwen3.8-27B-DFlash2
Controlled back-to-back A/B on Helga (4×3090 TP4, vLLM DFlash2 fork, n6, thinking-low, greedy, 3 prompt types × 2 reps):
| checkpoint | tok/s | accept len | Δ tok/s | Δ accept |
|---|---|---|---|---|
| this drafter (step 400) | 123.2 | 4.24 | +5.4% | +5.8% |
| stock z-lab | 116.9 | 4.01 | — | — |
Spec-off baseline on the same setup: 38.1 tok/s.
Quick start (vLLM)
vllm serve heretic-org/Qwen3.8-27B-heretic-ara \
--tensor-parallel-size 4 \
--reasoning-parser qwen3 \
--speculative-config '{"model": "alphakek/Qwen3.8-27B-heretic-ara-DFlash2", "method": "dflash", "num_speculative_tokens": 6}'
Requires vLLM with DFlash2 support (PR #52816).
License
Apache-2.0