Qwen3-8B Eagle3 Drafter (Qwen3 Architecture)
Eagle3 speculative decoding drafter for Qwen/Qwen3-8B using the Qwen3 draft architecture (--draft-arch qwen3).
The Qwen3 architecture uses additional q_norm/k_norm layers in attention, which stabilize training at higher learning rates compared to the default Llama draft architecture. See speculators#563 for the RFC and experimental results.
Benchmark Results
Evaluated with sampling params temp=0.6, top_p=0.95, top_k=20.
Acceptance lengths
| Use Case | k=1 | k=2 | k=3 | k=4 | k=5 | k=6 | k=7 |
|---|---|---|---|---|---|---|---|
| Coding (HumanEval, 164 samples) | 1.81 | 2.40 | 2.80 | 3.06 | 3.22 | 3.30 | 3.33 |
| Math Reasoning (gsm8k, 80 samples) | 1.83 | 2.50 | 2.95 | 3.28 | 3.45 | 3.61 | 3.68 |
| Text Summarization (CNN/DM, 80 samples) | 1.69 | 2.11 | 2.35 | 2.46 | 2.50 | 2.52 | 2.52 |
Comparison vs Llama-arch baseline
Trained on the same dataset with the same hyperparameters (except LR). Format: delta vs llamaarch-ckpt1.
| Use Case | k=1 | k=2 | k=3 | k=4 | k=5 | k=6 | k=7 |
|---|---|---|---|---|---|---|---|
| Coding | +0.02 | +0.05 | +0.09 | +0.18 | +0.22 | +0.24 | +0.23 |
| Math Reasoning | +0.02 | +0.08 | +0.08 | +0.16 | +0.15 | +0.19 | +0.23 |
| Text Summarization | +0.02 | +0.06 | +0.10 | +0.13 | +0.13 | +0.14 | +0.13 |
Outperforms the Llama-arch baseline by ~3-7% across all benchmarks and all k values, with the gap widening at higher k (math k=7: 3.68 vs 3.45, +6.7%).
Training
| Parameter | Value |
|---|---|
| Target model | Qwen/Qwen3-8B |
| Draft architecture | Qwen3 (--draft-arch qwen3) |
| Learning rate | 5e-4 |
| Epochs | 2 |
| Draft vocab size | 32000 |
| Sequence length | 8192 |
| Training mode | Online (hidden states generated on-the-fly) |
| Dataset | inference-optimization/Qwen3-8B-Regenerated-Collection (magpie + ultrachat subsets, 508k samples) |
| Hardware | 7x H200 (2 vLLM DP=2, 5 training FSDP) |
| Training library | speculators |
Usage
Requires vLLM with Qwen3 Eagle3 support (vllm#43132) and the architecture resolution fix.
vllm serve Qwen/Qwen3-8B \
--speculative-config '{
"model": "inference-optimization/Qwen3-8B-from-Qwen3-8B_regen-speculators.eagle3-qwen3arch-ckpt1",
"num_speculative_tokens": 3,
"method": "eagle3"
}'
Related
- speculators#563 — RFC: Support Qwen3 base architecture for Eagle3 & P-EAGLE
- Llama-arch baseline — Same setup with Llama draft architecture
- LR sweep results — Llama vs Qwen3 draft arch comparison across 5 learning rates