rdtand/Hy3-295B-A21B-prismaquant-gridbook-2.9bit-vllm

🤗 Hugging Face sourcetext-generationapache-2.0104B params21B activated106 GBsafetensors✓ 5 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo rdtand/Hy3-295B-A21B-prismaquant-gridbook-2.9bit-vllm ./model-folder
Needs a seeder →

Hy3-295B-A21B — PrismaQuant gridbook 2.9-bit (gridbook formats, single DGX Spark, native vLLM kernels)

A 295B-A21B MoE served on one 128 GB DGX Spark at 2.9 bits per parameter, through stock vLLM with an out-of-tree plugin — no forked runtime, no core patches. Weights are stored in the gridbook codebook formats (NVFP4-CB / FP8-CB: product-VQ over 8-weight vectors whose codebook values live on hardware grids, two-tier scale coding) and decoded inside custom CUDA/Triton kernels at matmul time, so prefill runs on native tensor cores instead of paying a dequantization tax. Formats + kernels: github.com/RobTand/gridbook. Produced by the PrismaQuant measured-allocation pipeline.

  • Base model: tencent/Hy3 (production release, Apache 2.0), 80 layers, 192 experts top-8 + 1 shared, GQA 8 KV heads.
  • Artifact: 5 safetensors shards, 105.7 GB total, 2.9 bpp over quantizable parameters. Joint-menu allocation, measured not hand-picked: codebook rungs compete with vanilla NVFP4/FP8 per Linear, and the measurement chose — 36 Linears landed on vanilla FP8; vanilla NVFP4 was offered and never won a unit. Full breakdown in Format allocation.
  • MTP included, CB-quantized at FP8-CB K44 (2.6 GB vs 7.5 GB BF16), serving as a k=1 speculative-decode drafter out of the box. The rung was selected by a throughput-optimal draft-quantization algorithm (derivation and fitted constants in allocation/mtp_rung_selection.json). Draft quantization can only affect acceptance rate — rejection sampling reproduces the target distribution exactly. A GGUF export cannot carry MTP at all.
  • Allocation provenance in allocation/ (per-Linear assignment, Pareto sweep, MTP rung selection).

Speed (measured on one DGX Spark, vLLM 0.23, this repo's serve.sh)

Metric This artifact GGUF IQ 2.8 bpp (same box)
Prefill (1.5k-token prompt) ~109 tok/s 42 tok/s
Decode (batch 1, base) 14.6 tok/s 17.8 tok/s
Decode (batch 1, prose, MTP spec on) 16.1 tok/s not possible

The prefill gap is the point of the formats: IQ/k-quant GGUF decodes weights on the fly in CUDA-core kernels; gridbook expands each layer transiently and runs native tensor-core GEMMs, then frees the tile. Base batch-1 decode still trails the GGUF k-quant/MMVQ path (the fp4-CB decode chain is compute-bound at GEMV shapes — measured and documented in the gridbook repo); MTP speculative decode closes most of that gap on natural text and becomes a straight multiplier once vLLM captures drafter CUDA graphs.

Format allocation

Exact on-disk accounting from the shipped shards + quant_config.json (CB totals include packed cb_qweight and fp8-rung weight_scale sidecars). A unit is one allocation target: for routed experts, one stacked tensor holding all 192 experts of one projection — two stacks per MoE layer; otherwise one Linear. "bits/w" is the rung's body bit-rate (NVFP4-CB two-tier: k/8 + 0.28125; FP8-CB: k/8 plus per-channel scale; vanilla FP8: 8).

Naming note: "FP8"/"NVFP4" in a CB format names the grid the codebook's VALUES live on, not the storage width. Storage is the k-bit vector index; the reconstructed values are exact e4m3 (or E2M1) numbers — which is why an expanded FP8-CB tile is a standard per-channel fp8 tensor that vLLM's stock fp8 tensor-core GEMM consumes directly. "FP8-CB K32" is a 4.0-bit weight encoding served through native fp8 kernels.

Format bits/w Role Units Size (GB)
NVFP4-CB K14 2.03 routed experts 18 stacks 8.281
NVFP4-CB K14 2.03 shared expert 25 0.040
NVFP4-CB K14 2.03 dense/attention 16 0.082
NVFP4-CB K16 2.28 routed experts 36 stacks 18.601
NVFP4-CB K16 2.28 shared expert 8 0.014
NVFP4-CB K16 2.28 dense/attention 3 0.012
NVFP4-CB K18 2.53 routed experts 38 stacks 21.786
NVFP4-CB K18 2.53 dense/attention 2 0.035
NVFP4-CB K20 2.78 routed experts 28 stacks 17.638
NVFP4-CB K20 2.78 dense/attention 6 0.029
FP8-CB K28 3.5+ routed experts 38 stacks 30.228
FP8-CB K28 3.5+ dense/attention 26 0.331
FP8-CB K32 4.0+ dense/attention 194 1.790
FP8-CB K32 4.0+ shared expert 133 0.419
FP8-CB K36 4.5+ dense/attention 33 0.326
FP8-CB K44 5.5+ dense/attention 23 0.289
FP8-CB K44 5.5+ shared expert 8 0.035
FP8-CB K44 5.5+ MTP experts + attention + shared 9 2.562
vanilla FP8 8 dense/attention 16 0.243
vanilla FP8 8 shared expert 20 0.126
BF16 16 embeddings 1 0.990
BF16 16 lm_head 1 0.990
BF16 16 shared-expert Linears (allocator-assigned) 43 0.541
BF16 16 norms (incl. q/k norms) 325 0.152
BF16 16 router gates 79 0.124
BF16 16 MTP glue (eh_proj, enorm/hnorm, norms) 9 0.069
F32 32 expert bias (routing) 80 <0.001
Total 105.732

Family rollup: NVFP4-CB 180 units / 66.5 GB · FP8-CB 464 units / 36.0 GB · vanilla FP8 36 units / 0.37 GB · BF16+F32 538 units / 2.9 GB.

Per-layer view of the 79 MoE layers (experts are format-uniform within a layer, mixed across layers — a vLLM fused-MoE serving constraint): 9 layers on NVFP4-CB K14 · 18 on K16 · 19 on K18 · 14 on K20 · 19 on FP8-CB K28. Layer 0 is dense; layer "80" is the MTP module (FP8-CB K44 throughout + BF16 glue).

No quality claims

A 295B cannot be KL-validated against its BF16 teacher on this hardware, so — as with our GGUF Hy3 releases — this card makes no quality claims. The validation ledger: vLLM loads and serves it; generation is coherent and factually/arithmetically correct on our smoke suite; packing is bit-exact against the reference codec (test-pinned); tool-calling behaviour was exercised end-to-end under the hardmode harness we run on every release. Quantization is calibrated (activation-weighted codebook encode; measured per-Linear format allocation on real calibration signals; at validatable scales — 27B/35B — the same formats measured −53 to −58% served KL against same-size conventional artifacts).

Serve this model

Plugin gridbook — an out-of-tree vLLM quantization plugin. Stock vLLM, no fork, no core patches.
GPU NVIDIA Blackwell, compute capability sm_120 / sm_121. Measured on GB10 / DGX Spark (sm_121). On older GPUs the plugin still loads but runs its Triton fallback kernels — correct, not fast, and not a production serving target.
Memory ~106 GB resident weights (including the CB-quantized MTP drafter) on a ~121 GB usable unified pool — one 128 GB DGX Spark. --gpu-memory-utilization 0.90 is the validated setting; higher values OOM'd under long-prefill activation spikes.
vLLM ≥ 0.23 with HYV3ForCausalLM (stock — no fork). Speed numbers on this card were measured on vLLM 0.23.
Toolchain CUDA toolkit with nvcc on PATH in the serving container — the plugin JIT-builds its kernels on first model load (~30 s, cached). nvcc 13.0 is the tested toolchain.
Parallelism Single GPU (tp=1). Tensor parallelism is not implemented in the plugin.

Requirements: vLLM >= 0.23 with HYV3ForCausalLM (stock), CUDA toolkit with nvcc (the plugin JIT-builds its kernels on first load), ~121 GB usable GPU/unified memory.

pip install gridbook
MODEL_DIR=. ./serving/serve.sh          # 0.0.0.0:8000, 12k ctx, fp8 KV, MTP on

The gridbook plugin (serving/gridbook/, Apache 2.0 — canonical home github.com/RobTand/gridbook) registers the gridbook quantization method through vLLM's standard plugin entry point and serves the codebook formats with:

  • fused activation-QDQ CUDA GEMV kernels for fp8-CB dense decode (double-buffered),
  • grouped MoE GEMV CUDA kernels (fp8-CB and fp4-CB two-tier v2) covering all routed (token, expert) pairs in one launch per projection,
  • transient tensor-core GEMM prefill (expand one layer's tile, matmul, free) — vanilla NVFP4/FP8 groups delegate to vLLM's stock compressed-tensors paths,
  • Triton fallbacks for every path (PRISMAQUANT_CB_DECODE=triton).

Memory: serve at --gpu-memory-utilization 0.90 (validated; higher settings OOM under long-prefill activation spikes on a 128 GB Spark). Context 12288 validated; fp8 KV recommended.

Attribution

gridbook formats + kernels and PrismaQuant allocation pipeline — Robert Tand (robert.tand@icloud.com). Base model by Tencent (Apache 2.0; see LICENSE).