rdtand/Hy3-295B-A21B-prismaquant-gridbook-2.9bit-vllm

🤗 Hugging Face 来源text-generationapache-2.0104B 参数激活 21B106 GBsafetensors✓ 5 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo rdtand/Hy3-295B-A21B-prismaquant-gridbook-2.9bit-vllm ./model-folder
需要做种者 →

Hy3-295B-A21B — PrismaQuant gridbook 2.9-bit (gridbook formats, single DGX Spark, native vLLM kernels)

A 295B-A21B MoE served on one 128 GB DGX Spark at 2.9 bits per parameter, through stock vLLM with an out-of-tree plugin — no forked runtime, no core patches. Weights are stored in the gridbook codebook formats (NVFP4-CB / FP8-CB: product-VQ over 8-weight vectors whose codebook values live on hardware grids, two-tier scale coding) and decoded inside custom CUDA/Triton kernels at matmul time, so prefill runs on native tensor cores instead of paying a dequantization tax. Formats + kernels: github.com/RobTand/gridbook. Produced by the PrismaQuant measured-allocation pipeline.

  • Base model: tencent/Hy3 (production release, Apache 2.0), 80 layers, 192 experts top-8 + 1 shared, GQA 8 KV heads.
  • Artifact: 5 safetensors shards, 105.7 GB total, 2.9 bpp over quantizable parameters. Joint-menu allocation, measured not hand-picked: codebook rungs compete with vanilla NVFP4/FP8 per Linear, and the measurement chose — 36 Linears landed on vanilla FP8; vanilla NVFP4 was offered and never won a unit. Full breakdown in Format allocation.
  • MTP included, CB-quantized at FP8-CB K44 (2.6 GB vs 7.5 GB BF16), serving as a k=1 speculative-decode drafter out of the box. The rung was selected by a throughput-optimal draft-quantization algorithm (derivation and fitted constants in allocation/mtp_rung_selection.json). Draft quantization can only affect acceptance rate — rejection sampling reproduces the target distribution exactly. A GGUF export cannot carry MTP at all.
  • Allocation provenance in allocation/ (per-Linear assignment, Pareto sweep, MTP rung selection).

Speed (measured on one DGX Spark, vLLM 0.23, this repo's serve.sh)

Metric This artifact GGUF IQ 2.8 bpp (same box)
Prefill (1.5k-token prompt) ~109 tok/s 42 tok/s
Decode (batch 1, base) 14.6 tok/s 17.8 tok/s
Decode (batch 1, prose, MTP spec on) 16.1 tok/s not possible

The prefill gap is the point of the formats: IQ/k-quant GGUF decodes weights on the fly in CUDA-core kernels; gridbook expands each layer transiently and runs native tensor-core GEMMs, then frees the tile. Base batch-1 decode still trails the GGUF k-quant/MMVQ path (the fp4-CB decode chain is compute-bound at GEMV shapes — measured and documented in the gridbook repo); MTP speculative decode closes most of that gap on natural text and becomes a straight multiplier once vLLM captures drafter CUDA graphs.

Format allocation

Exact on-disk accounting from the shipped shards + quant_config.json (CB totals include packed cb_qweight and fp8-rung weight_scale sidecars). A unit is one allocation target: for routed experts, one stacked tensor holding all 192 experts of one projection — two stacks per MoE layer; otherwise one Linear. "bits/w" is the rung's body bit-rate (NVFP4-CB two-tier: k/8 + 0.28125; FP8-CB: k/8 plus per-channel scale; vanilla FP8: 8).

Naming note: "FP8"/"NVFP4" in a CB format names the grid the codebook's VALUES live on, not the storage width. Storage is the k-bit vector index; the reconstructed values are exact e4m3 (or E2M1) numbers — which is why an expanded FP8-CB tile is a standard per-channel fp8 tensor that vLLM's stock fp8 tensor-core GEMM consumes directly. "FP8-CB K32" is a 4.0-bit weight encoding served through native fp8 kernels.

Format bits/w Role Units Size (GB)
NVFP4-CB K14 2.03 routed experts 18 stacks 8.281
NVFP4-CB K14 2.03 shared expert 25 0.040
NVFP4-CB K14 2.03 dense/attention 16 0.082
NVFP4-CB K16 2.28 routed experts 36 stacks 18.601
NVFP4-CB K16 2.28 shared expert 8 0.014
NVFP4-CB K16 2.28 dense/attention 3 0.012
NVFP4-CB K18 2.53 routed experts 38 stacks 21.786
NVFP4-CB K18 2.53 dense/attention 2 0.035
NVFP4-CB K20 2.78 routed experts 28 stacks 17.638
NVFP4-CB K20 2.78 dense/attention 6 0.029
FP8-CB K28 3.5+ routed experts 38 stacks 30.228
FP8-CB K28 3.5+ dense/attention 26 0.331
FP8-CB K32 4.0+ dense/attention 194 1.790
FP8-CB K32 4.0+ shared expert 133 0.419
FP8-CB K36 4.5+ dense/attention 33 0.326
FP8-CB K44 5.5+ dense/attention 23 0.289
FP8-CB K44 5.5+ shared expert 8 0.035
FP8-CB K44 5.5+ MTP experts + attention + shared 9 2.562
vanilla FP8 8 dense/attention 16 0.243
vanilla FP8 8 shared expert 20 0.126
BF16 16 embeddings 1 0.990
BF16 16 lm_head 1 0.990
BF16 16 shared-expert Linears (allocator-assigned) 43 0.541
BF16 16 norms (incl. q/k norms) 325 0.152
BF16 16 router gates 79 0.124
BF16 16 MTP glue (eh_proj, enorm/hnorm, norms) 9 0.069
F32 32 expert bias (routing) 80 <0.001
Total 105.732

Family rollup: NVFP4-CB 180 units / 66.5 GB · FP8-CB 464 units / 36.0 GB · vanilla FP8 36 units / 0.37 GB · BF16+F32 538 units / 2.9 GB.

Per-layer view of the 79 MoE layers (experts are format-uniform within a layer, mixed across layers — a vLLM fused-MoE serving constraint): 9 layers on NVFP4-CB K14 · 18 on K16 · 19 on K18 · 14 on K20 · 19 on FP8-CB K28. Layer 0 is dense; layer "80" is the MTP module (FP8-CB K44 throughout + BF16 glue).

No quality claims

A 295B cannot be KL-validated against its BF16 teacher on this hardware, so — as with our GGUF Hy3 releases — this card makes no quality claims. The validation ledger: vLLM loads and serves it; generation is coherent and factually/arithmetically correct on our smoke suite; packing is bit-exact against the reference codec (test-pinned); tool-calling behaviour was exercised end-to-end under the hardmode harness we run on every release. Quantization is calibrated (activation-weighted codebook encode; measured per-Linear format allocation on real calibration signals; at validatable scales — 27B/35B — the same formats measured −53 to −58% served KL against same-size conventional artifacts).

Serve this model

Plugin gridbook — an out-of-tree vLLM quantization plugin. Stock vLLM, no fork, no core patches.
GPU NVIDIA Blackwell, compute capability sm_120 / sm_121. Measured on GB10 / DGX Spark (sm_121). On older GPUs the plugin still loads but runs its Triton fallback kernels — correct, not fast, and not a production serving target.
Memory ~106 GB resident weights (including the CB-quantized MTP drafter) on a ~121 GB usable unified pool — one 128 GB DGX Spark. --gpu-memory-utilization 0.90 is the validated setting; higher values OOM'd under long-prefill activation spikes.
vLLM ≥ 0.23 with HYV3ForCausalLM (stock — no fork). Speed numbers on this card were measured on vLLM 0.23.
Toolchain CUDA toolkit with nvcc on PATH in the serving container — the plugin JIT-builds its kernels on first model load (~30 s, cached). nvcc 13.0 is the tested toolchain.
Parallelism Single GPU (tp=1). Tensor parallelism is not implemented in the plugin.

Requirements: vLLM >= 0.23 with HYV3ForCausalLM (stock), CUDA toolkit with nvcc (the plugin JIT-builds its kernels on first load), ~121 GB usable GPU/unified memory.

pip install gridbook
MODEL_DIR=. ./serving/serve.sh          # 0.0.0.0:8000, 12k ctx, fp8 KV, MTP on

The gridbook plugin (serving/gridbook/, Apache 2.0 — canonical home github.com/RobTand/gridbook) registers the gridbook quantization method through vLLM's standard plugin entry point and serves the codebook formats with:

  • fused activation-QDQ CUDA GEMV kernels for fp8-CB dense decode (double-buffered),
  • grouped MoE GEMV CUDA kernels (fp8-CB and fp4-CB two-tier v2) covering all routed (token, expert) pairs in one launch per projection,
  • transient tensor-core GEMM prefill (expand one layer's tile, matmul, free) — vanilla NVFP4/FP8 groups delegate to vLLM's stock compressed-tensors paths,
  • Triton fallbacks for every path (PRISMAQUANT_CB_DECODE=triton).

Memory: serve at --gpu-memory-utilization 0.90 (validated; higher settings OOM under long-prefill activation spikes on a 128 GB Spark). Context 12288 validated; fp8 KV recommended.

Attribution

gridbook formats + kernels and PrismaQuant allocation pipeline — Robert Tand (robert.tand@icloud.com). Base model by Tencent (Apache 2.0; see LICENSE).