inference-optimization/Qwen3.8-27B-DSpark-PerfectBlend-NVFP4-W4A4

🤗 Hugging Face 来源text-generationapache-2.02B 参数2.1 GBsafetensors✓ 2 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo inference-optimization/Qwen3.8-27B-DSpark-PerfectBlend-NVFP4-W4A4 ./model-folder
需要做种者 →

Qwen3.8-27B-DSpark-PerfectBlend-NVFP4-W4A4

A Static NVFP4 W4A4, target-calibrated DSpark drafter derived from RedHatAI/Qwen3.8-27B-speculator.dspark at revision 7f33c272e5da240978e0d55767abab8193d74b95. Pair it with Qwen/Qwen3.8-27B at revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0. This repository contains a drafter component, not a standalone chat model.

Quantization

Static NVFP4 W4A4 using 1,892 aligned target hidden-state calibration records and sequence cap 2,048.

This arm used a local proportional sample from shanjiaz/OpenPerfectBlend-Qwen38-27B-regenerated, derived from mlabonne/open-perfectblend; it is not the unmodified upstream PerfectBlend collection. The sample requested 2,048 examples; 1,892 aligned hidden-state records were available and used, with a sequence cap of 2,048. Quantization commands, manifests, source files, patches, and the exported checkpoint checksum are in provenance/quantization/.

Example serving command

vllm serve Qwen/Qwen3.8-27B \
  --spec-model inference-optimization/Qwen3.8-27B-DSpark-PerfectBlend-NVFP4-W4A4 \
  --spec-method dspark \
  --spec-tokens 8 \
  --kernel-config '{"linear_backend":"emulation"}'

For NVFP4, this example selects vLLM's emulation backend; it makes no claim of native H100 NVFP4 support.

Evaluation status

Evaluation is pending. No completed acceptance, speed, or quality results are included with this publication. The checkpoint has quantization provenance only; serving/runtime validation and the planned evaluation matrix have not completed.

Reproducibility

provenance/quantization/ includes the training and quantization commands, quantization manifest, calibration metadata where applicable, source scripts and patches, and the SHA-256 digest of the published drafter weights. Calibration prompt data is not redistributed.