ToBeStyled/Signal-3.8-27B-NVFP4-Blackwell-DFlash2-Ultra-V1.0

🤗 Hugging Face 来源text-generationapache-2.0激活 27B19 GBGGUF✓ 3 个校验和今天更新
需要做种者 →

Signal-3.8-27B-NVFP4-Blackwell-DFlash2-Ultra-V1.0

Terse answers, faster decode. An NVFP4 GGUF of agentionai/Signal-3.8-27B for one RTX 5090 (32 GB, sm_120) with llama.cpp CUDA and DFlash2 speculative decoding. 262K context.

Signal is a minimally invasive fine-tune of Qwen3.8-27B: only lm_head.weight differs from stock (head delta norm 4.1%), which is why it keeps the base model's body byte-identical — and why a base-trained drafter stays aligned with it. Quantized here in NVFP4 with high-precision heads.

What this package is

File Size
…-NVFP4.gguf 16.9 GB Model — NVFP4 backbone, Q8_0 lm_head + token_embd, MTP head included
…-draft-DFlash2-Q4_K_M.gguf 1.14 GB DFlash2 drafter — the one to use (measured below)
…-mmproj-BF16.gguf 0.93 GB Vision projector (optional)

Measured, on this box (RTX 5090 32 GB, 262K, DFlash2)

Everything below is interleaved A/B against the GAIN-based sibling package (ToBeStyled/Qwen3.8-27B-ColdFusion-GAIN-Blackwell-DFlash2-Ultra-V1.0), same codec (NVFP4), same file size (16.87 GB both sides), same drafter, same flags. Only interleaved comparisons are trusted — on this box, separate measurement windows drift by up to 10% with background load.

Structured and reasoning (temp 0.6, fixed seeds, GSM8K-50)

Config Score Tokens/problem tok/s s/problem DFlash2 acceptance
Signal-NVFP4, low 100 % (30/30) 302 107.1 — 5.93
Signal-NVFP4, medium 100 % (50/50) 363 157.6 2.30 5.78
GAIN, medium 96 % (48/50) 332 138.0 2.41 5.06

Read it as: at the same effort, Signal gives +14.2% decode via +14.2% acceptance and −4.6% time per problem, with equal-or-better accuracy. And reasoning_effort=low (which carries the only brevity instruction in the stock Qwen3.8 template; medium is a no-op there) cuts Signal to 302 tokens/problem at 100% held.

Agentic code task (2048 engine, 3 files, hidden acceptance tests, temp 0.7)

Success Turns Tokens Wall DFlash2 acceptance
Signal-NVFP4 2/2 2 6 803 41.8 s 6.33 (+4.2%)
GAIN 2/2 2 6 950 42.7 s 6.07

Both models solve the task in the same number of turns. Signal is marginally faster with a measurably better-aligned drafter.

Open conversation (15 general prompts, temp 0.7, deterministic)

GAIN Signal-NVFP4
total tokens 8 985 (FR) / 7 671 (EN) 12 235 / 10 534 (+36–37 %)

Honest caveat: on open chat, this build is more verbose than the GAIN package. Upstream reports −57% answer tokens against the stock model — against an already-terse finetune, the advantage disappears on conversational prompts and shows up only on structured/reasoning work. Use GAIN for chat, Signal for code and reasoning.

Run (llama.cpp CUDA, RTX 5090)

llama-server.exe \
  -m  ...-NVFP4.gguf \
  -md ...-draft-DFlash2-Q4_K_M.gguf \
  --spec-type draft-dflash --spec-draft-n-max 7 \
  -ngl all -np 1 -c 262144 -fa on \
  -ctk q8_0 -ctv q8_0 \
  --spec-draft-type-k f16 --spec-draft-type-v f16 \
  --chat-template-kwargs "{\"reasoning_effort\":\"low\"}" \
  --cache-ram 32768 --no-cache-idle-slots \
  --alias signal38-ultra:latest --reasoning-format auto
# add --mmproj ...-mmproj-BF16.gguf for image input (disables prompt-cache reuse)

Sampling — read this first. Use temperature 0.6 coding / 0.7 general, top_p 0.95, top_k 20. Do not use greedy (temperature 0): upstream documents degeneration on long generations at greedy, and one community report describes a repeat loop after 15–30 minutes on a from-scratch game build. For long builds also give the agent 64k+ context — a 30-minute build overflows 32k, and once the window shifts the model loses the start of its own code.

Reasoning effort: low is the recommended default here (fewest tokens at equal accuracy in our GSM8K sweeps). medium decodes slightly faster per token on math. xhigh is the template default; it was not re-measured for this build.

VRAM (measured, NVFP4 + DFlash2)

Config Slots Context per slot VRAM
Single agent, max context 1 262 144 29 447 MiB

Two-slot and f16-KV variants follow the GAIN package's table; the envelope is within 0.5 GB of it (identical file size).

Built from

  • Base: Qwen/Qwen3.8-27B
  • Finetune: agentionai/Signal-3.8-27B ("directness" self-distill; only lm_head.weight differs from stock)
  • NVFP4 quant: converted from BF16 with convert_hf_to_gguf.py @ f3f1a8f27, then llama-quantize with a per-tensor map over a 200-chunk imatrix (NVFP4 backbone, Q8_0 heads). The recipes that put lm_head or token_embd at 6-bit or below visibly damage this model — both were measured and rejected.
  • Drafter: z-lab/Qwen3.8-27B-DFlash2 (block-diffusion, lossless). Speculative decoding is lossless (rejection sampling), so the drafter never changes outputs, only speed.

Notes

  • Thinking model — set max_tokens ≥ 4096 on hard tasks.
  • 262K context works; vision pairs with the included mmproj-BF16.gguf.
  • CUDA graphs stay ON (measured +33% decode on this box in the sibling package).
  • Ollama: see the GAIN sibling repo's Modelfile for the template (same one works here).
  • Only interleaved A/B runs are trusted here; see the sibling repo for the methodology and the harnesses (ab-run-07.ps1, eval-gsm8k-07.py, agentic.py).