sakamakismile/Agents-A1-4B-NVFP4

🤗 Hugging Face sourceapache-2.03B params3.9 GBsafetensors✓ 3 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo sakamakismile/Agents-A1-4B-NVFP4 ./model-folder
Needs a seeder →

Agents-A1-4B-NVFP4

NVFP4 (W4A4, compressed-tensors) quantization of InternScience/Agents-A1-4B — a 4B agentic model on the Qwen3.5 dense architecture (Qwen3_5ForConditionalGeneration: hybrid DeltaNet linear attention + full attention every 4 layers, Qwen3-VL vision tower, 262K context).

  • Size: 4.21 GB (from ~8 GB bf16), incl. 0.24 GB bf16 MTP draft head
  • Bonus — grafted MTP: the source checkpoint declares mtp_num_hidden_layers: 1 but ships no mtp.* weights. This repo grafts the 15-tensor bf16 MTP draft head from the base model Qwen/Qwen3.5-4B (identical text-config dims), enabling speculative decoding: 63–66% draft acceptance, ~+20% single-stream decode (measured, see below)
  • Scheme: NVFP4 W4A4, group size 16, via llm-compressor 0.11.0 / compressed-tensors 0.16.0
  • Kept bf16: lm_head, vision tower (model.visual*), DeltaNet conv1d
  • Calibration: 32 samples × 8192 seq, neuralmagic/calibration (pure-CPU calibration — no GPU used in the bake)

Serve with vLLM

vllm serve sakamakismile/Agents-A1-4B-NVFP4

Quantization is auto-detected — no --quantization flag needed. Requires a GPU with FP4 support (SM120 Blackwell) for the NVFP4 kernels.

With speculative decoding (grafted MTP draft, ~+20% single-stream):

vllm serve sakamakismile/Agents-A1-4B-NVFP4 \
  --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":3}'

Measured throughput — single GPU (vLLM 0.22.0, RTX PRO 2000 Blackwell SM120, KV fp8, 512-tok decode)

Config single-stream (c1) aggregate 8-way (c8) vs base
base (no MTP) 73.6 t/s 492.9 t/s —
MTP n=3 (grafted) 88.6 t/s 474.5 t/s +20.4% c1 / −3.7% c8

SpecDecoding metrics (vLLM): mean acceptance length ~2.9, per-position acceptance 0.83 / 0.66 / 0.48, avg draft acceptance 63–66% — remarkable for a draft head grafted from the pre-fine-tune base model across InternScience's 3-stage agentic distillation. As usual, MTP is a latency win (single-stream / low concurrency); at saturation the extra draft forward costs slightly more than it saves.

Speculative decoding is lossless — outputs are identical to base decoding.

Recipe

Baked (pure-CPU, llm-compressor) with the same validated recipe as Qwen3.6-27B-MTP-pi-tune-NVFP4 and ThinkingCap-Qwen3.6-27B-NVFP4 (same qwen3_5 architecture family).