Hy3-295B-A21B — PrismaQuant gridbook 2.9-bit (gridbook formats, single DGX Spark, native vLLM kernels)
A 295B-A21B MoE served on one 128 GB DGX Spark at 2.9 bits per parameter, through stock vLLM with an out-of-tree plugin — no forked runtime, no core patches. Weights are stored in the gridbook codebook formats (NVFP4-CB / FP8-CB: product-VQ over 8-weight vectors whose codebook values live on hardware grids, two-tier scale coding) and decoded inside custom CUDA/Triton kernels at matmul time, so prefill runs on native tensor cores instead of paying a dequantization tax. Formats + kernels: github.com/RobTand/gridbook. Produced by the PrismaQuant measured-allocation pipeline.
- Base model: tencent/Hy3 (production release, Apache 2.0), 80 layers, 192 experts top-8 + 1 shared, GQA 8 KV heads.
- Artifact: 5 safetensors shards, 105.7 GB total, 2.9 bpp over quantizable parameters. Joint-menu allocation, measured not hand-picked: codebook rungs compete with vanilla NVFP4/FP8 per Linear, and the measurement chose — 36 Linears landed on vanilla FP8; vanilla NVFP4 was offered and never won a unit. Full breakdown in Format allocation.
- MTP included, CB-quantized at FP8-CB K44 (2.6 GB vs 7.5 GB BF16), serving
as a k=1 speculative-decode drafter out of the box. The rung was selected by
a throughput-optimal draft-quantization algorithm (derivation and fitted
constants in
allocation/mtp_rung_selection.json). Draft quantization can only affect acceptance rate — rejection sampling reproduces the target distribution exactly. A GGUF export cannot carry MTP at all. - Allocation provenance in
allocation/(per-Linear assignment, Pareto sweep, MTP rung selection).
Speed (measured on one DGX Spark, vLLM 0.23, this repo's serve.sh)
| Metric | This artifact | GGUF IQ 2.8 bpp (same box) |
|---|---|---|
| Prefill (1.5k-token prompt) | ~109 tok/s | 42 tok/s |
| Decode (batch 1, base) | 14.6 tok/s | 17.8 tok/s |
| Decode (batch 1, prose, MTP spec on) | 16.1 tok/s | not possible |
The prefill gap is the point of the formats: IQ/k-quant GGUF decodes weights on the fly in CUDA-core kernels; gridbook expands each layer transiently and runs native tensor-core GEMMs, then frees the tile. Base batch-1 decode still trails the GGUF k-quant/MMVQ path (the fp4-CB decode chain is compute-bound at GEMV shapes — measured and documented in the gridbook repo); MTP speculative decode closes most of that gap on natural text and becomes a straight multiplier once vLLM captures drafter CUDA graphs.
Format allocation
Exact on-disk accounting from the shipped shards + quant_config.json (CB
totals include packed cb_qweight and fp8-rung weight_scale sidecars). A
unit is one allocation target: for routed experts, one stacked tensor
holding all 192 experts of one projection — two stacks per MoE layer; otherwise
one Linear. "bits/w" is the rung's body bit-rate (NVFP4-CB two-tier:
k/8 + 0.28125; FP8-CB: k/8 plus per-channel scale; vanilla FP8: 8).
Naming note: "FP8"/"NVFP4" in a CB format names the grid the codebook's VALUES live on, not the storage width. Storage is the k-bit vector index; the reconstructed values are exact e4m3 (or E2M1) numbers — which is why an expanded FP8-CB tile is a standard per-channel fp8 tensor that vLLM's stock fp8 tensor-core GEMM consumes directly. "FP8-CB K32" is a 4.0-bit weight encoding served through native fp8 kernels.
| Format | bits/w | Role | Units | Size (GB) |
|---|---|---|---|---|
| NVFP4-CB K14 | 2.03 | routed experts | 18 stacks | 8.281 |
| NVFP4-CB K14 | 2.03 | shared expert | 25 | 0.040 |
| NVFP4-CB K14 | 2.03 | dense/attention | 16 | 0.082 |
| NVFP4-CB K16 | 2.28 | routed experts | 36 stacks | 18.601 |
| NVFP4-CB K16 | 2.28 | shared expert | 8 | 0.014 |
| NVFP4-CB K16 | 2.28 | dense/attention | 3 | 0.012 |
| NVFP4-CB K18 | 2.53 | routed experts | 38 stacks | 21.786 |
| NVFP4-CB K18 | 2.53 | dense/attention | 2 | 0.035 |
| NVFP4-CB K20 | 2.78 | routed experts | 28 stacks | 17.638 |
| NVFP4-CB K20 | 2.78 | dense/attention | 6 | 0.029 |
| FP8-CB K28 | 3.5+ | routed experts | 38 stacks | 30.228 |
| FP8-CB K28 | 3.5+ | dense/attention | 26 | 0.331 |
| FP8-CB K32 | 4.0+ | dense/attention | 194 | 1.790 |
| FP8-CB K32 | 4.0+ | shared expert | 133 | 0.419 |
| FP8-CB K36 | 4.5+ | dense/attention | 33 | 0.326 |
| FP8-CB K44 | 5.5+ | dense/attention | 23 | 0.289 |
| FP8-CB K44 | 5.5+ | shared expert | 8 | 0.035 |
| FP8-CB K44 | 5.5+ | MTP experts + attention + shared | 9 | 2.562 |
| vanilla FP8 | 8 | dense/attention | 16 | 0.243 |
| vanilla FP8 | 8 | shared expert | 20 | 0.126 |
| BF16 | 16 | embeddings | 1 | 0.990 |
| BF16 | 16 | lm_head | 1 | 0.990 |
| BF16 | 16 | shared-expert Linears (allocator-assigned) | 43 | 0.541 |
| BF16 | 16 | norms (incl. q/k norms) | 325 | 0.152 |
| BF16 | 16 | router gates | 79 | 0.124 |
| BF16 | 16 | MTP glue (eh_proj, enorm/hnorm, norms) | 9 | 0.069 |
| F32 | 32 | expert bias (routing) | 80 | <0.001 |
| Total | 105.732 |
Family rollup: NVFP4-CB 180 units / 66.5 GB · FP8-CB 464 units / 36.0 GB · vanilla FP8 36 units / 0.37 GB · BF16+F32 538 units / 2.9 GB.
Per-layer view of the 79 MoE layers (experts are format-uniform within a layer, mixed across layers — a vLLM fused-MoE serving constraint): 9 layers on NVFP4-CB K14 · 18 on K16 · 19 on K18 · 14 on K20 · 19 on FP8-CB K28. Layer 0 is dense; layer "80" is the MTP module (FP8-CB K44 throughout + BF16 glue).
No quality claims
A 295B cannot be KL-validated against its BF16 teacher on this hardware, so — as with our GGUF Hy3 releases — this card makes no quality claims. The validation ledger: vLLM loads and serves it; generation is coherent and factually/arithmetically correct on our smoke suite; packing is bit-exact against the reference codec (test-pinned); tool-calling behaviour was exercised end-to-end under the hardmode harness we run on every release. Quantization is calibrated (activation-weighted codebook encode; measured per-Linear format allocation on real calibration signals; at validatable scales — 27B/35B — the same formats measured −53 to −58% served KL against same-size conventional artifacts).
Serve this model
| Plugin | gridbook — an out-of-tree vLLM quantization plugin. Stock vLLM, no fork, no core patches. |
| GPU | NVIDIA Blackwell, compute capability sm_120 / sm_121. Measured on GB10 / DGX Spark (sm_121). On older GPUs the plugin still loads but runs its Triton fallback kernels — correct, not fast, and not a production serving target. |
| Memory | ~106 GB resident weights (including the CB-quantized MTP drafter) on a ~121 GB usable unified pool — one 128 GB DGX Spark. --gpu-memory-utilization 0.90 is the validated setting; higher values OOM'd under long-prefill activation spikes. |
| vLLM | ≥ 0.23 with HYV3ForCausalLM (stock — no fork). Speed numbers on this card were measured on vLLM 0.23. |
| Toolchain | CUDA toolkit with nvcc on PATH in the serving container — the plugin JIT-builds its kernels on first model load (~30 s, cached). nvcc 13.0 is the tested toolchain. |
| Parallelism | Single GPU (tp=1). Tensor parallelism is not implemented in the plugin. |
Requirements: vLLM >= 0.23 with HYV3ForCausalLM (stock), CUDA toolkit with
nvcc (the plugin JIT-builds its kernels on first load), ~121 GB usable
GPU/unified memory.
pip install gridbook
MODEL_DIR=. ./serving/serve.sh # 0.0.0.0:8000, 12k ctx, fp8 KV, MTP on
The gridbook plugin (serving/gridbook/, Apache 2.0 — canonical home
github.com/RobTand/gridbook) registers the gridbook quantization method
through vLLM's standard plugin entry point and serves the codebook formats
with:
- fused activation-QDQ CUDA GEMV kernels for fp8-CB dense decode (double-buffered),
- grouped MoE GEMV CUDA kernels (fp8-CB and fp4-CB two-tier v2) covering all routed (token, expert) pairs in one launch per projection,
- transient tensor-core GEMM prefill (expand one layer's tile, matmul, free) — vanilla NVFP4/FP8 groups delegate to vLLM's stock compressed-tensors paths,
- Triton fallbacks for every path (
PRISMAQUANT_CB_DECODE=triton).
Memory: serve at --gpu-memory-utilization 0.90 (validated; higher settings
OOM under long-prefill activation spikes on a 128 GB Spark). Context 12288
validated; fp8 KV recommended.
Attribution
gridbook formats + kernels and PrismaQuant allocation pipeline —
Robert Tand (robert.tand@icloud.com).
Base model by Tencent (Apache 2.0; see LICENSE).