inference-optimization/Qwen3-8B-DFlash-Gauss-NVFP4-W4A4

🤗 Hugging Face 来源text-generationapache-2.01.2B 参数1.3 GBsafetensors✓ 2 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo inference-optimization/Qwen3-8B-DFlash-Gauss-NVFP4-W4A4 ./model-folder
需要做种者 →

Gauss4-NVFP4-W4A4 — NVFP4 W4A4 with random Gaussian calibration

Gauss4 is a quantized DFlash drafter for the Qwen3-8B target, derived from RedHatAI/Qwen3-8B-speculator.dflash. This repository contains the drafter component; it is not a standalone chat model.

Variant

  • Quantization: NVFP4 W4A4 (weight group size 16; local-dynamic activations).
  • Calibration: Synthetic random Gaussian calibration: 2,027 samples, sequence length 2,048, seed 0. Weight observer: nvfp4_expanded_mse. This is a calibration control, not real-data calibration.
  • Calibration seed: 0.
  • Quantization settings and source revisions: quant_run_manifest.json.

For NVFP4 evaluation on H100, the serving backend used W4A4 emulation; this artifact does not claim native Blackwell NVFP4 serving performance.

Use with vLLM

Pair this drafter with the Qwen3-8B target and a DFlash-capable vLLM build:

vllm serve Qwen/Qwen3-8B \
  --spec-model inference-optimization/Qwen3-8B-DFlash-Gauss4-NVFP4-W4A4 \
  --spec-tokens 7 \
  --spec-method dflash

config.py provides the custom drafter configuration. The experiment's serving command and runtime patch are in provenance/evaluation/.

Reproducibility

The manifests are included at the repository root. provenance/ contains the source drafter's captured train_command.txt, a quantization command explicitly marked as reconstructed, the quantizer and calibration source snapshot, the vLLM command and patch, both target and drafter checkpoint hashes, and the nine per-subset evaluation commands. The selected seed-0 checkpoint is the same checkpoint used in the 2026-09-24 all-subset evaluation.

The PerfectBlend preparation, prompts, and hidden-state cache remain local because the prepared prompts are not redistributable. The cache sample counts and content hashes are recorded in calibration_manifest.json; no prompts or hidden-state tensors are uploaded.

The source drafter lists Apache-2.0 licensing on its Hugging Face model card.