brandonmusic/glm-5.3-flash-tr3-fp8

🤗 Hugging Face sourcemit193 GBsafetensors✓ 294 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo brandonmusic/glm-5.3-flash-tr3-fp8 ./model-folder
Needs a seeder →

glm-5.3-flash-tr3-fp8

TrellisMX codec-v2 with Trellis MTP experts and the corrected native FP8 runtime.

TrellisMX codec-v2 quantization of GLM-5.3-Flash, including all 42 main routed layers and all 288 MTP45 experts at K4. This complete package includes its native carrier, tokenizer, configuration, and 172 TP4 sidecars. No separate base-weight download is required. Loading requires the included custom TrellisMX runtime.

Compressed Trellis indices decode into E4M3 operands with UE8M0/K32 scales for native FP8 Tensor Core expert execution. The checkpoint is not stored as a full 8-bit weight copy; other modules retain the carrier's precision. The tr3-fp8 repository name is the owner's chosen name for this TrellisMX release.

KLD

Native FP8-KV serving-system comparison

The completed local native run measured 0.028330843421 mean teacher-to-student KLD, compared with 0.030099944949 for the saved EXL3/TR3 quantization of the same model. Both use the same 128 windows, exact token sequences, teacher logits, native FP8-KV target-only forced-decode protocol, and full-vocabulary CPU FP64 scorer. The saved TR3 run was reused without rerunning inference.

Native target-only measurement Mean KLD, nats
TrellisMX codec-v2, TP4/DCP4, FP8 KV 0.028330843421
Saved EXL3/TR3, TP4/DCP4, FP8 KV 0.030099944949

On this panel, the observed native TrellisMX serving-system mean was 5.88% lower than the saved native EXL3/TR3 mean. TrellisMX was lower on 48 of 128 paired windows. This reports observed system behavior; the quantizations, backends, image revisions, and non-routed carrier precision policies differ, so the difference cannot be attributed solely to codec-v2. It is one prepared run on previously used conditional-fit development data, not an untouched holdout or a statistical equivalence claim.

Routed-weight storage budgets also differ: TrellisMX has a 4.25-bit-per-weight payload (four-bit indices plus eight scale bits per 32 weights). The saved TR3 uses four-bit trellises plus per-channel overhead, without those block scales; see its pinned configuration and the EXL3 storage implementation. This is not an equal-bit-budget comparison.

The lower mean is concentrated: TR3 is lower on the other 80 windows and on the median paired window. Code-agentic windows account for most of the summed difference, while legal windows average about 3% higher KLD for TrellisMX. The domain breakdown and raw paired results retain these differences. No confidence interval or significance claim is made.

This full-panel KLD was measured on the earlier unlimited-promotion runtime. The final prefill-corrected image has not had the full panel remeasured; its single-window checks did not establish numerical equivalence. The KLD report preserves the distinct image identities and check results.

The score covers 261,888 true-decode positions, excluding each window's one-token-prefill row. MTP was disabled; draft quality, speculative acceptance, and high-concurrency quality were not measured. See the native report and limitations, raw summary, and independent audit.

Separate B200 reference result

The earlier B200 reference runner measured 0.027881131803 mean KLD over the same 128-window panel. It used a BF16 non-routed backbone, unquantized KV, full-window Transformers 5.17 causal forwards, and FP32 expert matmuls. This is a separate reference-path measurement; it is not the native FP8-KV result above and is not substituted for it. Its numerical 7.37% difference from historical native EXL3/TR3 cannot by itself establish matched-serving equivalence or superiority. The reference run also does not measure MTP draft quality or speculative acceptance. See the B200 reference report.

Earlier-runtime behavioral results

Each profile used 30 measured requests at concurrency 10, Max reasoning, temperature 1.0, and top-p 0.95 on four RTX PRO 6000 Blackwell GPUs. The original prompts and scorer were retained. TP4/DCP4, FP8 KV, and MTP1 were enabled on image sha256:8e45776f38283d1e0b2ab0422e3f9a01bf11b5d55889b6cff8497cdeea1a8343, with 48 maximum sequences, forced M16 tiles, and the extended 192-row direct cutoff. All 90 requests finished normally, with zero request errors and zero truncations. These are preserved results from that earlier runtime, not quality measurements of the new MTP3/compact-capture image.

Profile Automated result Reviewed result
Hotel Lights 29/30 30/30 correct blue-light answers; one final-line formatting/extraction miss
Estonia 29/30 29/30; one response ultimately answered Latvia
LAVD 30/30 passes 28 exact, 2 near, under the original tolerance

Hotel's missed response explicitly answers 48, then ends with 48 + 52 = 100; the automated final-line scorer extracts 100. Its automated score remains unchanged. LAVD's two near answers are 68 tickets / 42.75 hours, versus the exact 72 / 46; both fall within the existing tolerance of four in each quantity. These are repeated fixed-prompt results, not a broad model-quality evaluation.

See complete results, machine-readable summary, and immutable raw responses and score reviews.

Reproduce the runtime

The Docker Compose and serve.sh use the published image verdictai/trellismx:glm53-codecv2-prefill-20260927, pinned by digest to sha256:6240b05f889c2096cfcf62a8873de43ba09ac4f745d2e31f5464496fb9d76916. The registry manifest was verified through unauthenticated Docker Hub access and matches the measured image. The complete build recipe and source overlays are included.

From the complete checkpoint directory, launch:

export CHECKPOINT_ROOT=/absolute/path/to/glm-5.3-flash-tr3-fp8
chmod +x serve.sh benchmark.sh
./serve.sh

The complete checkpoint can be downloaded with hf download brandonmusic/glm-5.3-flash-tr3-fp8 --revision codecv2-prefill-20260927 --local-dir /absolute/path/to/glm-5.3-flash-tr3-fp8. Root config.json is included for the Hub's standard download accounting; Hugging Face controls the counter and cached transfers may not increment it.

The corrected configuration uses TP4/DCP4, FP8 KV, MTP3, automatic small-row tiles (B12X_P8_RP2_DIRECT_TILE_M=0), the original 16-row direct cutoff, and 128 maximum sequences. Sampling defaults remain temperature 1.0 / top-p 0.95. No diagnostic profiler is enabled by the supplied Compose. LOCAL-SERVING.md contains the build and benchmark commands.

The speed fix binds compact MTP prefill outputs to the active request count inside each graph capture. Previously, a C1 graph could retain work for the configured 128-request capacity; the corrected graph performs the draft work for its actual request count. The earlier RP2 addressing, state-padding, and FP32 BF16-SUM fixes remain in place. Expert matmuls retain native FP8 Tensor Core execution; checkpoint weights are unchanged.

The saved default-sampling decode results measured 189.2 tokens/s at C1/0K and 573.4 aggregate tokens/s at C8/0K. Prefill measured 9,520 tokens/s at nominal 32K and 9,372 at nominal 128K (128,880 actual tokens, 13.751-second client TTFT, one sample). 33 decode conditions completed; the owner ended further decode measurement, leaving 15 conditions not run. See the saved decode results, actual concurrency, unrun conditions, and raw evidence. KLD and behavioral results above retain their original measured runtime identities.

The prefill correction adds VLLM_PYNCCL_BF16_SUM_FP32_MAX_BYTES=8388608. Eligible BF16 SUM reductions up to and including 8 MiB retain FP32 communication/accumulation; larger reductions use native BF16 NCCL. This avoids doubling the collective payload for large prefill batches while retaining the corrected small-reduction path. The limit is based on original tensor bytes, not a prefill/decode label, so large mixed batches also change arithmetic. Batched tokens remain 8,192. This is a runtime change; checkpoint weights are unchanged.

The launch restores the original dispatch policy: direct RP2 through 16 token rows, grouped execution above 16. The installed safe layout extension still supports up to 192 direct rows if explicitly selected, but it is not the current launch default. With MTP3, C1 verification fits in four rows, while C8 can reach 32 rows and takes the grouped path. Above 16 rows, the original grouped path uses its single-term activation carrier; this is a precision-policy change from the extended two-term direct path, not a claim of numerical equivalence to it. Throughput does not validate answer quality at a concurrency. See the bounded dispatch source review.

Checkpoint contents and encoding

The measured checkpoint revision is 3730bd667c83405c6f0e6835c6fb7f27e28658b0. At that revision, all 499 repository files total 185,123,620,576 bytes (185.124 GB / 172.410 GiB); this includes evidence and scripts. The 172 routed sidecars total 165,721,576,928 bytes. The complete carrier directory is 19,388,363,261 bytes. New release documentation and result files add a small amount to repository download size.

The main codec uses routed plus all-token Hessians with beta 0.5, damping 0.3, rounded power-of-two UE8M0 scales, and K4 throughout. Physical routed-weight payload is 4.25 bits per weight. Carrier tensor bytes and configuration are preserved; only replaced routed tensors are omitted from its index. See carrier receipt, manifest, and verification report.

All 288 MTP45 experts are Trellis K4 in four TP4 sidecars totaling 3,853,990,880 bytes. Its 25 non-expert tensors are preserved exactly. Encoding used the existing BF16 NextN captures: 64 fit windows with 2,047 shifted-token rows, all routed rows, every fourth all-token row, beta 0.5, and damping 0.3. MTP uses deterministic signed-unit coupled vectors with CPU seed 530045; main-model vectors are unchanged. All 864 MTP projection encodes passed packed/decode closure, and all 288 expert streams reassemble exactly across TP4 sidecars. See MTP design and MTP verification.

Source identities

  • BF16: zai-org/GLM-5.3-Flash-BF16@a6c167b62691b2bac901344b65cb651a70f53e43
  • Teacher logits: brandonmusic/GLM-5.3-Flash-BF16-Teacher-Logits@95f4fdd94bf29989db2e0d1054e4931f55edb6aa
  • r27: brandonmusic/GLM-5.3-Flash-TrellisMX-MXFP8@ea7e0ea310242dc0ce9a5faa9069eab6ebdbae9e
  • Native carrier: local-inference-lab/GLM-5.3-Flash-NVFP4@520de24eabf507659eaef7c70f14fd584527facc
  • Historical EXL3/TR3 reference: brandonmusic/GLM-5.3-Flash-tr3-4bpw@4bccf1bcf9dcd3357933a8ad91193de675331012
  • Supplied encoder bundle SHA256: eeee96b2d09df5da04d44abf04d7dc93a97e9fa8e3b04978de91af0f9c232d73

Additional benchmark runners

benchmarks/README.md contains the owner's inference benchmark, pinned datasets, and commands for LAVD, Estonia, Hotel, needle, MMLU-Pro, and GSM8K. This release's completed behavioral suite covers Hotel, Estonia, and LAVD only; it does not report GSM8K, MMLU-Pro, or needle scores.

License

The base-model weights and native carrier retain Z.AI's MIT License. Bundled software retains its own notices and licenses, including the runtime Apache-2.0 text and third-party notices. Existing licenses and permissions remain in effect. No new restriction is imposed on rights previously granted for the model or its dependencies.

An independent evidence audit rechecked all 90 saved answers, raw hashes, request settings, reported means, and the 128-window KLD identities.