local-inference-lab/Qwen3.8-27B-NVFP4-QAD

🤗 Hugging Face 来源image-text-to-textapache-2.019.2B 参数24 GBsafetensors✓ 68 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo local-inference-lab/Qwen3.8-27B-NVFP4-QAD ./model-folder
需要做种者 →

*** WORK IN PROGRESS ***

Qwen3.8-27B-NVFP4-QAD

A mixed NVFP4/MXFP8 quantization-aware distillation of Qwen3.8-27B, exported at training step 7,000.

The student learns from the original BF16 teacher with quantized MLP weights in its forward pass. Distillation updates the MLP weights, text normalization weights and BF16 student LM head. This is a trained distillation checkpoint, not a post-training conversion of the original weights.

Precision

Component Representation
MLP gate, up and down projections NVFP4, 16-element blocks
Gated delta network projections MXFP8, 32-element blocks; frozen
Full-attention query, key, value and output projections Original BF16; frozen
Student LM head Trained FP32 master rounded to BF16
Text normalization weights Trained FP32 masters
Token embeddings, GDN convolutions and dynamics Original BF16; frozen
Vision encoder and remaining source tensors Unchanged

Packed NVFP4 and MXFP8 weights reconstruct to the BF16 weight values used during training. The tokenizer, chat template, generation configuration and multimodal processors are retained from the base model.

Activation scales

The 192 MLP input scales are copied exactly from Qwen3.8-27B-QAD-E1, whose step-4,779 weights were calibrated on 390,497,191 raw-text and chat tokens. That calibration selected the pooled p99.999 token-row maximum from exact BF16 histograms. These scales were not recalibrated at step 7,000.

Each dense layer has equal gate/up scales in separate tensors and an independent down-projection scale. Serving uses calibrated NVFP4 MLP activations and dynamic MXFP8 GDN activations, adding activation quantization beyond the BF16 training forward.

Format

Hugging Face safetensors with ModelOpt mixed-precision metadata. The runtime must support the base architecture, NVFP4 dense linears, MXFP8 linears and the BF16 full-attention/head exclusions. Packed reconstruction, serialized tensors and copied input scales are checked. Serving quality and performance have not been evaluated for this export.

License

Apache 2.0, following the base model.