inference-optimization/Qwen3-8B-from-Qwen3-8B_regen-speculators.eagle3-qwen3arch-ckpt1

🤗 Hugging Face sourceapache-2.01B params2.0 GBsafetensors✓ 1 checksumupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo inference-optimization/Qwen3-8B-from-Qwen3-8B_regen-speculators.eagle3-qwen3arch-ckpt1 ./model-folder
Needs a seeder →

Qwen3-8B Eagle3 Drafter (Qwen3 Architecture)

Eagle3 speculative decoding drafter for Qwen/Qwen3-8B using the Qwen3 draft architecture (--draft-arch qwen3).

The Qwen3 architecture uses additional q_norm/k_norm layers in attention, which stabilize training at higher learning rates compared to the default Llama draft architecture. See speculators#563 for the RFC and experimental results.

Benchmark Results

Evaluated with sampling params temp=0.6, top_p=0.95, top_k=20.

Acceptance lengths

Use Case k=1 k=2 k=3 k=4 k=5 k=6 k=7
Coding (HumanEval, 164 samples) 1.81 2.40 2.80 3.06 3.22 3.30 3.33
Math Reasoning (gsm8k, 80 samples) 1.83 2.50 2.95 3.28 3.45 3.61 3.68
Text Summarization (CNN/DM, 80 samples) 1.69 2.11 2.35 2.46 2.50 2.52 2.52

Comparison vs Llama-arch baseline

Trained on the same dataset with the same hyperparameters (except LR). Format: delta vs llamaarch-ckpt1.

Use Case k=1 k=2 k=3 k=4 k=5 k=6 k=7
Coding +0.02 +0.05 +0.09 +0.18 +0.22 +0.24 +0.23
Math Reasoning +0.02 +0.08 +0.08 +0.16 +0.15 +0.19 +0.23
Text Summarization +0.02 +0.06 +0.10 +0.13 +0.13 +0.14 +0.13

Outperforms the Llama-arch baseline by ~3-7% across all benchmarks and all k values, with the gap widening at higher k (math k=7: 3.68 vs 3.45, +6.7%).

Training

Parameter Value
Target model Qwen/Qwen3-8B
Draft architecture Qwen3 (--draft-arch qwen3)
Learning rate 5e-4
Epochs 2
Draft vocab size 32000
Sequence length 8192
Training mode Online (hidden states generated on-the-fly)
Dataset inference-optimization/Qwen3-8B-Regenerated-Collection (magpie + ultrachat subsets, 508k samples)
Hardware 7x H200 (2 vLLM DP=2, 5 training FSDP)
Training library speculators

Usage

Requires vLLM with Qwen3 Eagle3 support (vllm#43132) and the architecture resolution fix.

vllm serve Qwen/Qwen3-8B \
  --speculative-config '{
    "model": "inference-optimization/Qwen3-8B-from-Qwen3-8B_regen-speculators.eagle3-qwen3arch-ckpt1",
    "num_speculative_tokens": 3,
    "method": "eagle3"
  }'

Related