sakamakismile/KAT-Coder-V2.5-Dev-NVFP4

🤗 Hugging Face 来源apache-2.020.4B 参数22 GBsafetensors✓ 2 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo sakamakismile/KAT-Coder-V2.5-Dev-NVFP4 ./model-folder
需要做种者 →

KAT-Coder-V2.5-Dev-NVFP4

NVFP4 (W4A4, compressed-tensors) quantization of Kwaipilot/KAT-Coder-V2.5-Dev — the 35B-A3B agentic-coding MoE (SWE-bench Verified 69.4% upstream).

70GB bf16 → 21.9GB. Runs on a single 24GB Blackwell card, or 2× 16GB.

Quantized by Lna-Lab (@Tono_Ken3).

Measured (12x RTX PRO 2000 Blackwell 16GB, 60W power cap each)

  • TP=2 (2 GPUs): ~122 tok/s single-stream
  • Sanity: lookahead-bias trading question answered with the correct shift(1) fix; clean O(n) implementations with tests.

Serving

vllm serve <this-repo> --tensor-parallel-size 2 --max-model-len 32768 \
  --gpu-memory-utilization 0.92
# NVFP4 is auto-detected; no --quantization flag needed.
# On no-P2P multi-GPU boxes add: NCCL_P2P_DISABLE=1, --disable-custom-all-reduce

Recipe notes

  • Arch qwen3_5_moe quantizes cleanly with llm-compressor (targets=Linear, scheme NVFP4) with these ignores: lm_head, re:.*conv1d.* (DeltaNet conv), re:.*mlp.gate$ and re:.*shared_expert_gate$ (MoE routers), re:.*mtp.*.
  • The open-weight release ships without vision and without MTP tensors (we checked; nothing to graft).
  • Calibration: 32 samples x 8192 tokens (neuralmagic/calibration), single GPU, ~4 minutes total wall.
  • Newer transformers removed GraniteMoeParallelExperts which llm-compressor still imports — a dummy class injected before import satisfies it safely for non-Granite models.

W4A4 is stable on this architecture (5th model family we've confirmed: Nex-N2, Holo, Ornith, AgentWorld, KAT).