vcruz305/StepFun-5-Preview-GGUF

🤗 Hugging Face 来源text-generationapache-2.0494 GBGGUF✓ 56 个校验和今天更新
帮助这个模型通过 Pirate Face 分发

模型卡、文件列表和校验和已在此收录。如果你持有文件并有权分享,可以提交种子,让其他人从节点下载。

获取来源文件需要 Hugging Face 批准。
为此模型做种

Step-5-Preview-GGUF

GGUF quants of Step-5-Preview — a 604B-class sparse MoE (measured 598.04 B parameters across 1,548 tensors), 352 routed experts with top-8 routing, ~16B active, 95 layers (92 text + 3 MTP), 128,896 vocab, 1,048,576 context.

This repo is private and is not a general release. The quant packs carry the text trunk — 598.04 B of the checkpoint's 602.97 B parameters. The remaining 4.9325 B ship alongside them as GGUF companion files; see Companion files. They sit outside the packs because llama.cpp's step35 implements no runtime for them, not because they were dropped.

Layout

One folder per quant type, with the quant type in every file name — the convention unsloth uses, so the Hub and every downstream tool can identify each pack without being told.

Folder Shards Size bpw Fits
IQ1_S/ 25 111.32 GiB 1.599 one 128 GB GB10 / DGX Spark-class box
Q2_K/ 25 (20 uploaded) 202.45 GiB 2.908 two GB10-class boxes

Point llama.cpp at shard 1 only — it loads the remaining shards of the split itself.

# one box
hf download vcruz305/StepFun-5-Preview-GGUF --include "IQ1_S/*" --local-dir .
llama-cli -m IQ1_S/Step-5-Preview-IQ1_S-00001-of-00025.gguf -ngl 99 -c 512 \
  -p "The capital of France is" -n 64 --temp 0

# two boxes
hf download vcruz305/StepFun-5-Preview-GGUF --include "Q2_K/*" --local-dir .
llama-cli -m Q2_K/Step-5-Preview-Q2_K-00001-of-00025.gguf -ngl 99 -c 512 \
  -p "The capital of France is" -n 64 --temp 0

Requires a llama.cpp with the step35 architecture (any recent build).

Quants

Every pack below is quantized from the same rope-fixed BF16 conversion with the same importance matrix, so they differ only in bit width — not in coverage or metadata. Split into 25 shards each; download one pack with --include "<QUANT>/*".

quant size bpw shards fits
IQ1_S/ 111.32 GiB 1.599 25/25 shards 1 box
Q2_K/ 202.45 GiB 2.908 25/25 shards 2 boxes
Q3_K_M/ 265.29 GiB 3.811 25/25 shards 3 boxes

Metadata contract (important)

A converter that gets these wrong produces a file that loads perfectly and cannot speak.

  • step35.rope.freq_base = 1e7 (full-attention layers) and step35.rope.freq_base_swa = 1e4 (sliding layers). Writing one value into both is the defect this repo's packs were rebuilt to fix.
  • Full-attention layers rotate 96 of 192 head dims (step35.cpp halves n_rot_full); no override needed.
  • Sliding window 512. Routed experts: sigmoid routing, router bias applied after the activation for selection only, weights normalised over the selected set and scaled by 3.0.
  • metadata-rope.txt records the rope KV pairs as written, for verification.

Known limitations

  • IQ1_S is literate but degraded — grammatical, on-topic English, but reasoning depth is measurably gone (see the perplexity above). Usable-on-one-box, not a substitute for the BF16 checkpoint or a 2-bit quant on two boxes.
  • Routed experts in IQ1_S run at 1.5625 bpw, a ~10:1 compression of 584B parameters. Lossy regardless of calibration; the next type up (IQ2_XXS) needs ~150 GB for the experts alone and therefore two boxes.
  • All 161 sparse_indexer_* and 23 ssmax_s tensors are absent: llama.cpp has no runtime for this checkpoint's compressed sparse attention, so the 23 full-attention layers run dense. At ≤512-token contexts the indexer's top-512 selection is a no-op, but its constant softmax scale (0.08496 against llama.cpp's 1/sqrt(192) = 0.07217) is not applied by stock llama.cpp.
  • Importance matrix: v2, domain-mixed.

Companion files — the rest of the model

The packs hold the text trunk. The checkpoint's remaining 4.9325 B parameters ship alongside as GGUF sidecars. Tensor bytes are copied verbatim; the norms and biases that are FP32 in the checkpoint stay FP32, so nothing is re-rounded:

  • mmproj-BF16.gguf — 3.69 GiB, 667 tensors, 1.9810 B params. The perception encoder (vision_model.* + vit_large_projector.weight): ViT-class, 47 layers, width 1536, image 728, patch 14, quick_gelu. Its architecture key is perception_encoder, deliberately not clip — llama.cpp has no loader for this encoder, and a file that reads as clip would fail as a shape mismatch instead of reporting an unknown architecture.
  • MTP/mtp-Step-5-Preview-BF16.gguf — 4.69 GiB, 51 tensors, 2.5159 B params. The three MTP layers 92–94, each eh_proj + enorm + hnorm plus a full block.
  • sparse-indexer/indexer-Step-5-Preview-BF16.gguf — 0.81 GiB, 184 tensors, 0.4356 B params. The compressed sparse-attention indexer (sparse_indexer_*) and its ssmax_s scale for the 23 full-attention layers.

On the folder names: the Hub describes a GGUF repo using the first gguf path in lexical order, so no sidecar may sort before IQ1_S/. sparse-indexer/ and MTP/ sort after the quant folders, and mmproj-BF16.gguf carries the prefix the Hub recognises as a projector — that combination is what keeps this card describing the model instead of a sidecar. Renaming one of them to sort earlier (e.g. INDEXER/) makes the repo report the wrong parameter count.

All 902 tensors across the three were compared byte-for-byte against the checkpoint — 0 mismatches — and each file's sha256 was re-checked against the Hub copy after upload.

They are not inside the quant packs because none of the three has a runtime in llama.cpp's step35: a pack carrying them would still not execute them. A runtime that loads the model in HF format (vLLM / TensorRT-LLM) is not bound by that and can consume all of them.

Not published here

  • Q4_K_M for a 4-box setup (~340 GiB) — queued.