Step-5-Preview-GGUF
GGUF quants of Step-5-Preview — a 604B-class sparse MoE (measured 598.04 B parameters across 1,548 tensors), 352 routed experts with top-8 routing, ~16B active, 95 layers (92 text + 3 MTP), 128,896 vocab, 1,048,576 context.
This repo is private and is not a general release. The quant packs carry the text
trunk — 598.04 B of the checkpoint's 602.97 B parameters. The remaining
4.9325 B ship alongside them as GGUF companion files; see
Companion files. They sit outside the packs
because llama.cpp's step35 implements no runtime for them, not because they were
dropped.
Layout
One folder per quant type, with the quant type in every file name — the convention unsloth uses, so the Hub and every downstream tool can identify each pack without being told.
| Folder | Shards | Size | bpw | Fits |
|---|---|---|---|---|
IQ1_S/ |
25 | 111.32 GiB | 1.599 | one 128 GB GB10 / DGX Spark-class box |
Q2_K/ |
25 (20 uploaded) | 202.45 GiB | 2.908 | two GB10-class boxes |
Point llama.cpp at shard 1 only — it loads the remaining shards of the split itself.
# one box
hf download vcruz305/StepFun-5-Preview-GGUF --include "IQ1_S/*" --local-dir .
llama-cli -m IQ1_S/Step-5-Preview-IQ1_S-00001-of-00025.gguf -ngl 99 -c 512 \
-p "The capital of France is" -n 64 --temp 0
# two boxes
hf download vcruz305/StepFun-5-Preview-GGUF --include "Q2_K/*" --local-dir .
llama-cli -m Q2_K/Step-5-Preview-Q2_K-00001-of-00025.gguf -ngl 99 -c 512 \
-p "The capital of France is" -n 64 --temp 0
Requires a llama.cpp with the step35 architecture (any recent build).
Quants
Every pack below is quantized from the same rope-fixed BF16 conversion with the same
importance matrix, so they differ only in bit width — not in coverage or metadata.
Split into 25 shards each; download one pack with --include "<QUANT>/*".
| quant | size | bpw | shards | fits |
|---|---|---|---|---|
IQ1_S/ |
111.32 GiB | 1.599 | 25/25 shards | 1 box |
Q2_K/ |
202.45 GiB | 2.908 | 25/25 shards | 2 boxes |
Q3_K_M/ |
265.29 GiB | 3.811 | 25/25 shards | 3 boxes |
Metadata contract (important)
A converter that gets these wrong produces a file that loads perfectly and cannot speak.
step35.rope.freq_base = 1e7(full-attention layers) andstep35.rope.freq_base_swa = 1e4(sliding layers). Writing one value into both is the defect this repo's packs were rebuilt to fix.- Full-attention layers rotate 96 of 192 head dims (
step35.cpphalvesn_rot_full); no override needed. - Sliding window 512. Routed experts: sigmoid routing, router bias applied after the activation for selection only, weights normalised over the selected set and scaled by 3.0.
metadata-rope.txtrecords the rope KV pairs as written, for verification.
Known limitations
IQ1_Sis literate but degraded — grammatical, on-topic English, but reasoning depth is measurably gone (see the perplexity above). Usable-on-one-box, not a substitute for the BF16 checkpoint or a 2-bit quant on two boxes.- Routed experts in
IQ1_Srun at 1.5625 bpw, a ~10:1 compression of 584B parameters. Lossy regardless of calibration; the next type up (IQ2_XXS) needs ~150 GB for the experts alone and therefore two boxes. - All 161
sparse_indexer_*and 23ssmax_stensors are absent: llama.cpp has no runtime for this checkpoint's compressed sparse attention, so the 23 full-attention layers run dense. At ≤512-token contexts the indexer's top-512 selection is a no-op, but its constant softmax scale (0.08496 against llama.cpp's1/sqrt(192)= 0.07217) is not applied by stock llama.cpp. - Importance matrix: v2, domain-mixed.
Companion files — the rest of the model
The packs hold the text trunk. The checkpoint's remaining 4.9325 B parameters ship alongside as GGUF sidecars. Tensor bytes are copied verbatim; the norms and biases that are FP32 in the checkpoint stay FP32, so nothing is re-rounded:
mmproj-BF16.gguf— 3.69 GiB, 667 tensors, 1.9810 B params. The perception encoder (vision_model.*+vit_large_projector.weight): ViT-class, 47 layers, width 1536, image 728, patch 14,quick_gelu. Its architecture key isperception_encoder, deliberately notclip— llama.cpp has no loader for this encoder, and a file that reads asclipwould fail as a shape mismatch instead of reporting an unknown architecture.MTP/mtp-Step-5-Preview-BF16.gguf— 4.69 GiB, 51 tensors, 2.5159 B params. The three MTP layers 92–94, eacheh_proj+enorm+hnormplus a full block.sparse-indexer/indexer-Step-5-Preview-BF16.gguf— 0.81 GiB, 184 tensors, 0.4356 B params. The compressed sparse-attention indexer (sparse_indexer_*) and itsssmax_sscale for the 23 full-attention layers.
On the folder names: the Hub describes a GGUF repo using the first gguf path in
lexical order, so no sidecar may sort before IQ1_S/. sparse-indexer/ and MTP/ sort
after the quant folders, and mmproj-BF16.gguf carries the prefix the Hub recognises as a
projector — that combination is what keeps this card describing the model instead of a
sidecar. Renaming one of them to sort earlier (e.g. INDEXER/) makes the repo report the
wrong parameter count.
All 902 tensors across the three were compared byte-for-byte against the checkpoint — 0 mismatches — and each file's sha256 was re-checked against the Hub copy after upload.
They are not inside the quant packs because none of the three has a runtime in llama.cpp's
step35: a pack carrying them would still not execute them. A runtime that loads the model
in HF format (vLLM / TensorRT-LLM) is not bound by that and can consume all of them.
Not published here
Q4_K_Mfor a 4-box setup (~340 GiB) — queued.