rdtand/Gemma4-31B-IT-PrismaQuant-6bit-vllm

🤗 Hugging Face 来源text-generationapache-2.022.2B 参数26 GBsafetensors✓ 7 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo rdtand/Gemma4-31B-IT-PrismaQuant-6bit-vllm ./model-folder
需要做种者 →

Gemma 4 31B IT — PrismaQuant 6-bit (vLLM)

A PrismaQuant 6.0 bits-per-weight mixed-native quantization of google/gemma-4-31b-it, exported in the compressed-tensors format for vLLM / FlashInfer serving on NVIDIA Blackwell (sm_120/sm_121) and Hopper.

This is the higher-fidelity sibling of rdtand/Gemma4-31B-IT-PrismaQuant-5.5bit-vllm: the language-model body is allocated at the 6.0-bpp Pareto knee of the allocator's predicted-Δloss-vs-bits curve (one step up from the 5.5-bit build), trading ~3 GB of extra size for measurably-closer-to-BF16 behavior (see Quality below). The 5.5-bit build remains available for a smaller footprint.

This is a vLLM-targeted checkpoint, not a vanilla Transformers model.

Contents & recipe

Component Format
Language-model body — 234 Linears NVFP4 (W4A4, group-16)
Language-model body — 135 Linears FP8 E4M3 (W8A8, per-channel)
Language-model body — 41 Linears BF16 (passthrough)
Norms / embeddings / lm_head / buffers BF16
Vision tower (355 tensors) BF16 passthrough
  • 6.0 bpp on the language-model body (410 Linears total).
  • ~27.2 GB on disk. The vision tower is carried in BF16 (the body is what is quantized); text generation is the validated path.
  • No audio tower.

Quality

Measured as KL divergence vs the BF16 reference (google/gemma-4-31b-it) on WikiText-2, teacher-forced with the BOS token, per-position top-20, fully deterministic. Lower KL / higher next-token agreement = closer to the original model.

Metric (vs BF16) 5.5-bit build This 6-bit build
KL-vs-BF16 (confident positions) ↓ 1.93 1.47 (−24%)
Next-token top-1 agreement ↑ 62.5% 68.4% (+5.9 pp)

The 6-bit build is consistently closer to BF16 than the 5.5-bit build on every axis measured. (Absolute KL is top-K-truncated, so the meaningful quantity is the relative gap between builds against the same reference. Plain perplexity does not separate these builds on this heavily instruction-tuned model — both sit near the BF16 value — which is why closeness-to-BF16 is reported instead.)

What is PrismaQuant?

PrismaQuant is a Fisher-weighted, mixed-precision quantization toolkit. Instead of forcing the whole model into one dtype, it predicts the loss penalty of quantizing each Linear independently — Δloss ≈ 0.5 · H_trace · MSE_W, where H_trace is the Linear's Fisher diagonal trace from a calibration probe and MSE_W is the format-specific reconstruction error — then solves for the per-Linear format menu ({NVFP4, FP8_E4M3, BF16}) that minimizes total predicted Δloss at the target bits-per-weight. NVFP4 Linears additionally go through act-aware GPTQ + closed-form scale-sweep passes, each gated "improve or keep RTN."

Usage (vLLM)

vllm serve rdtand/Gemma4-31B-IT-PrismaQuant-6bit-vllm \
  --quantization compressed-tensors --trust-remote-code

Requires a vLLM build with compressed-tensors NVFP4 + FP8 support (FlashInfer CUTLASS kernels on Blackwell; FP8 on Hopper+). Serve without speculative decoding if you intend to read prompt logprobs.

Attribution

Quantization by PrismaQuant. Contact: robert.tand@icloud.com. Base model © Google, under the Gemma 4 license.