inference-optimization/Qwen3-8B-DFlash-Gauss-NVFP4-W4A4

🤗 Hugging Face sourcetext-generationapache-2.01.2B params1.3 GBsafetensors✓ 2 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo inference-optimization/Qwen3-8B-DFlash-Gauss-NVFP4-W4A4 ./model-folder
Needs a seeder →

Gauss4-NVFP4-W4A4 — NVFP4 W4A4 with random Gaussian calibration

Gauss4 is a quantized DFlash drafter for the Qwen3-8B target, derived from RedHatAI/Qwen3-8B-speculator.dflash. This repository contains the drafter component; it is not a standalone chat model.

Variant

  • Quantization: NVFP4 W4A4 (weight group size 16; local-dynamic activations).
  • Calibration: Synthetic random Gaussian calibration: 2,027 samples, sequence length 2,048, seed 0. Weight observer: nvfp4_expanded_mse. This is a calibration control, not real-data calibration.
  • Calibration seed: 0.
  • Quantization settings and source revisions: quant_run_manifest.json.

For NVFP4 evaluation on H100, the serving backend used W4A4 emulation; this artifact does not claim native Blackwell NVFP4 serving performance.

Use with vLLM

Pair this drafter with the Qwen3-8B target and a DFlash-capable vLLM build:

vllm serve Qwen/Qwen3-8B \
  --spec-model inference-optimization/Qwen3-8B-DFlash-Gauss4-NVFP4-W4A4 \
  --spec-tokens 7 \
  --spec-method dflash

config.py provides the custom drafter configuration. The experiment's serving command and runtime patch are in provenance/evaluation/.

Reproducibility

The manifests are included at the repository root. provenance/ contains the source drafter's captured train_command.txt, a quantization command explicitly marked as reconstructed, the quantizer and calibration source snapshot, the vLLM command and patch, both target and drafter checkpoint hashes, and the nine per-subset evaluation commands. The selected seed-0 checkpoint is the same checkpoint used in the 2026-09-24 all-subset evaluation.

The PerfectBlend preparation, prompts, and hidden-state cache remain local because the prepared prompts are not redistributable. The cache sample counts and content hashes are recorded in calibration_manifest.json; no prompts or hidden-state tensors are uploaded.

The source drafter lists Apache-2.0 licensing on its Hugging Face model card.