peculiar-ragdoll/Nail-Qwen3.6-35B-A3B-GGUF-MTP

🤗 Hugging Face 来源image-text-to-textapache-2.0激活 3B263 GBGGUF✓ 18 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo peculiar-ragdoll/Nail-Qwen3.6-35B-A3B-GGUF-MTP ./model-folder
需要做种者 →

Straight to the point

Nail dominates Qwen3.6-27b and its ThinkingCap fine-tune on time-to-answer while matching their accuracy on reasoning and agentic software engineering ability, and wins on multi-turn conversation quality even when the 27b models use a higher quantization.

Say goodbye to overthinking, tool call failures, amnesic loops, and fluffy outputs.

Uncensored — Occult Nail 1.0 answers what Nail declines; use responsibly.

These are Unsloth's MTP GGUF quants of Qwen3.6-35B-A3B with our improved chat template and force-appended system prompt baked in — the identical recipe to the standard Nail GGUF, only this build keeps the multi-token-prediction (MTP / nextn) tensors that the standard quants drop. Runtimes that support MTP speculative decoding can use them for faster generation. llama.cpp applies the template automatically, nothing needed from you.

The benchmarks below were run on the non-MTP models. MTP adds a draft head for speculative decoding — accuracy is unchanged either way. Unsloth reports ~1.5–2× faster inference in vLLM/SGLang. In llama.cpp it needs a recent source build (b10362+) run with --spec-type draft-mtp; we measured ~1.35× on code generation tuned (--spec-draft-n-max 5 --spec-draft-p-min 0.6), ~1.25× at defaults (M2 Ultra, Q4). Current llama.cpp releases don't ship MTP yet, so there it's parity with the standard GGUF.

If you have a Mac, you can use the MLX build.

The numbers

All benchmarks on Mac Studio M2 Ultra 64 GB, oMLX, 8bit KV cache, MLX quants.

Reasoning and knowledge Multi turn conversations, capped at 6 turns Autonomous software engineering in Pi coding agent

Best conversation score, fastest to a working fix, fastest to a correct answer — against dense models that take 2-5x times as long, at accuracy parity with all of them. Four benchmarks, n=3 seeds, one Mac Studio M2 Ultra. Nail beats Opus4.8 (medium) 3-0 on first-attempt solves in a real SWE Live repo, no fluffing about. We used Pi for local models, Claude Code for Opus.

Which one. Many independent tasks, or you'd rather solve 3 problems in the time Qwen3.6-27b solves one → Nail. One long agentic session that has to stay coherent inside a single context → Dagger-27b, its dense sibling at the same ~21 GB RAM footprint.

Deployment math. The Nail GGUF and about 180k tokens of context in 8-bit KV should fit comfortably on a 24GB RAM card, which means it probably runs 120 tok/s on a single RTX4090 or RX7900 without partial offloading.

Memory footprint in multi-turn conversations to the context ceiling

Nail with full context and unquantized KV cache fits and runs on a 32GB unified RAM Mac, because it uses less RAM per token in context even at full 16bit KV precision.

While Nail runs all the way to 92 turns and 262k tokens context on that machine, Dagger taps out at 42 turns and 73k tokens.

Constrained hardware makes Nail the marathon winner AND the sprint winner, without sacrificing quality over Dagger or ThinkingCap.

Dagger wins for users with a lot of RAM, and a very long task, who are not in a hurry.

The other axis: context

Long form stamina is context ceiling ÷ tokens-per-question:

Seconds-per-correct is a sprint metric. In a long session that keeps every turn's thinking in context, the binding resource isn't time — it's context space, and your ceiling is context ÷ tokens per question.

Nail declares the same 262,144-token native context as Dagger, so the ceiling is not where it loses — the fill rate is. On GPQA-Diamond Nail spends 5,777 tokens per question against Dagger's 2,380:

  • Hard questions chained inside the native context: Dagger 110 · ThinkingCap 63 · Nail 45 · stock Qwen 24.
  • Conversation turns inside 100k tokens, on a harness that keeps each turn's thinking (Claw-Eval per-turn think+answer): Dagger 57 · ThinkingCap 44 · Nail 35.

Nobody sat through a 110-step session — it matters for autonomous thinking, agentic work, and conversations you keep coming back to.

Nail's verbose thinking is free in seconds and expensive in context. It still beats stock Qwen on both, but here Dagger wins by 2.4×. Nail for throughput, Dagger for endurance — the two axes have opposite winners, and neither model answers both.

The secret sauce

The system prompt ships inside the chat template and applies on every call. If you send your own system prompt, yours is emitted first and Nail's is appended after it:

You are Nail-35b-a3b, a variant of Qwen3.6-35B-A3B. Answer directly, after thinking. Lead with the answer, then only what it needs to be correct and usable.
Never: open with preamble or pleasantries; restate the question; add filler transitions; hedge with niceties; or repeat a point you've already made.
Always: keep essential steps, caveats, uncertainties, and specifics — never drop correctness or a needed warning for brevity. Keep the final answer lean. Use the least structure that conveys it (plain prose when short; lists or code only when they earn their place). If genuinely uncertain, say so and explain why — never omit uncertainty for the sake of brevity.
If a user request is genuinely ambiguous, ask a sharp question, don't guess.

Byte-identical to Dagger's, apart from the name.

The prompt is always on and lives in the template, not the API — disabling it means editing chat_template.jinja.

Froggeric's chat template that we used as a base implements many tricks that drastically improves multi-turn agentic workflows with tool calling.

The composition of our battle tested prompt with the improved template is what makes Nail so effective.

What the numbers are, and aren't

Every measurement we publish was taken on the MLX build, not on this file. They are not the same artifact: MLX Unsloth-Dynamic-4bit and GGUF UD-Q4_K_XL are different quantization schemes at a similar size, and llama.cpp and oMLX are different runtimes. On the one head-to-head we did run — the UD-Q4_K_XL file, llama-bench on an M2 Ultra, 2k prompt — llama.cpp was 1.6× faster on prefill and level-to-slightly-slower on decode than oMLX on the MLX build. The Q5 and Q6 files are unmeasured; expect slower decode roughly in proportion to their size.

So: treat the plate as evidence about the recipe — a 3.4B-active MoE with this template and this prompt — not as a measurement of this file. If you want numbers for this exact GGUF, we haven't taken them.

Use

Let llama.cpp fetch it — pass a :quant tag from the table below. The tag is required: this repo has no Q4_K_M, so a bare -hf with no tag falls back to the wrong file (the first one in the repo — a BF16 shard). The mmproj rides along in the manifest, so vision works from the same tag — no second download.

# text — auto-downloads to llama.cpp's own cache
llama-server   -hf peculiar-ragdoll/Nail-Qwen3.6-35B-A3B-GGUF-MTP:Q4_K_XL -ngl 99   # or llama-cli
# vision — same tag; the mmproj is pulled automatically
llama-mtmd-cli -hf peculiar-ragdoll/Nail-Qwen3.6-35B-A3B-GGUF-MTP:Q4_K_XL -ngl 99 --image photo.jpg

Swap :Q4_K_XL for any quant in the table (:IQ4_XS, :Q6_K_XL, …) — every quantized file is tag-addressable. The one exception is the sharded BF16, which has no single tag; grab its two shards with the explicit download below. Prefer to keep the files yourself? Download explicitly, then point -m at the local path:

hf download peculiar-ragdoll/Nail-Qwen3.6-35B-A3B-GGUF-MTP \
  --include "*MTP-UD-Q4_K_XL.gguf" "mmproj-F16.gguf" --local-dir Nail-MTP   # a bare download pulls ~223 GB
llama-cli      -m Nail-MTP/Nail-Qwen3.6-35B-A3B-MTP-UD-Q4_K_XL.gguf -ngl 99                          # text
llama-mtmd-cli -m Nail-MTP/Nail-Qwen3.6-35B-A3B-MTP-UD-Q4_K_XL.gguf \
               --mmproj Nail-MTP/mmproj-F16.gguf -ngl 99 --image photo.jpg                           # vision

Which file

file size notes
Nail-Qwen3.6-35B-A3B-MTP-UD-IQ3_XXS.gguf 14.1 GB smallest — the 16 GB VRAM entry (MTP files run a touch larger than non-MTP, so this is the one that leaves room for context)
Nail-Qwen3.6-35B-A3B-MTP-UD-IQ3_S.gguf 15.4 GB a step up; tight on a 16 GB card
Nail-Qwen3.6-35B-A3B-MTP-UD-IQ4_XS.gguf 18.2 GB the 24 GB pick — near-Q4_K_S quality (imatrix non-linear 4-bit) at ~3 GB less, leaving real room for long context + 8-bit KV and a display. Best all-round choice on a 24 GB card
Nail-Qwen3.6-35B-A3B-MTP-UD-Q4_K_S.gguf 21.4 GB 24 GB, max quality — the highest-fidelity quant that still fits 24 GB, but headroom is tight: prefer it headless or at modest context, otherwise take IQ4_XS above
Nail-Qwen3.6-35B-A3B-MTP-UD-Q4_K_XL.gguf 22.9 GB start here — the quant the head-to-head below was run on (non-MTP), and the size-match to a 27B at 6-bit
Nail-Qwen3.6-35B-A3B-MTP-UD-Q5_K_XL.gguf 27.2 GB more bits if you have the headroom
Nail-Qwen3.6-35B-A3B-MTP-UD-Q6_K_XL.gguf 32.6 GB most bits. On a 64 GB machine the MoE's cheap context makes this affordable where a dense 27B at the same quality would not be
Nail-Qwen3.6-35B-A3B-MTP-UD-Q8_K_XL.gguf 39.1 GB near-lossless 8-bit — the smallest quant with no measurable quality gap, and a big step under full BF16 on memory
Nail-Qwen3.6-35B-A3B-MTP-BF16-0000{1,2}-of-00002.gguf ~71 GB (2 shards) full precision — Unsloth's BF16 MTP weights, unquantized, with the Nail template. The maximum-fidelity build; wants a 96 GB+ machine. Download both shards; llama.cpp loads them as one

The eight quantized files are Unsloth Dynamic quants; the last is Unsloth's unquantized BF16. All carry the identical embedded template, and all use the same mmproj-F16.gguf for vision — you only need one copy of it. All are -hf tag addressable, e.g. llama-cli -hf peculiar-ragdoll/Nail-Qwen3.6-35B-A3B-GGUF-MTP:MTP-UD-Q6_K_XL.

Your own system prompt is preserved — the terseness directives are appended after it, so a harness with a large prompt of its own still gets its instructions honored. Override the concision directives explicitly and you have plain Qwen3.6-35B-A3B with a fixed template.

Recommended sampling

temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0

No repetition or presence penalties. Concision comes from the prompt; penalizing tokens distorts thinking in ways we haven't tested.

Limitations

  • The numbers above are not from this file, but from MLX. See above. This is the honest caveat.
  • The prompt is not removable through the API. It lives in the GGUF's embedded chat template. Use Unsloth's original MTP GGUFs if you need unmodified behavior.
  • Vision needs mmproj-F16.gguf and llama-mtmd-cli. Plain llama-cli is text-only.
  • We did not quantize this. Quality is entirely Unsloth's UD recipe; we changed one metadata string. Any quantization loss is theirs to characterize, and they document it better than we could.

Credits

  • Qwen at Alibaba — the Qwen3.6-35B-A3B base model.
  • Unsloth — the UD-Q4_K_XL quantization this repo redistributes.
  • froggeric — the fixed Qwen chat template.
  • llama.cpp — the runtime.

Citation

@misc{Nail-35B-A3B-GGUF-MTP,
  title  = {Nail-Qwen3.6-35B-A3B-GGUF-MTP},
  author = {Saga Ishtardottir},
  year   = {2026},
  url    = {https://huggingface.co/peculiar-ragdoll/Nail-Qwen3.6-35B-A3B-GGUF-MTP},
  note   = {Unsloth's MTP UD quants of Qwen3.6-35B-A3B with a fixed chat template and an always-on terseness prompt; MTP tensors preserved}
}

License

Apache-2.0, inherited from Qwen3.6-35B-A3B.