Signal-3.8-27B-NVFP4-Blackwell-DFlash2-Ultra-V1.0
Terse answers, faster decode. An NVFP4 GGUF of agentionai/Signal-3.8-27B for one RTX 5090
(32 GB, sm_120) with llama.cpp CUDA and DFlash2 speculative decoding. 262K context.
Signal is a minimally invasive fine-tune of Qwen3.8-27B: only lm_head.weight differs from stock
(head delta norm 4.1%), which is why it keeps the base model's body byte-identical — and why a
base-trained drafter stays aligned with it. Quantized here in NVFP4 with high-precision heads.
What this package is
| File | Size | |
|---|---|---|
…-NVFP4.gguf |
16.9 GB | Model — NVFP4 backbone, Q8_0 lm_head + token_embd, MTP head included |
…-draft-DFlash2-Q4_K_M.gguf |
1.14 GB | DFlash2 drafter — the one to use (measured below) |
…-mmproj-BF16.gguf |
0.93 GB | Vision projector (optional) |
Measured, on this box (RTX 5090 32 GB, 262K, DFlash2)
Everything below is interleaved A/B against the GAIN-based sibling package
(ToBeStyled/Qwen3.8-27B-ColdFusion-GAIN-Blackwell-DFlash2-Ultra-V1.0), same
codec (NVFP4), same file size (16.87 GB both sides), same drafter, same flags.
Only interleaved comparisons are trusted — on this box, separate measurement
windows drift by up to 10% with background load.
Structured and reasoning (temp 0.6, fixed seeds, GSM8K-50)
| Config | Score | Tokens/problem | tok/s | s/problem | DFlash2 acceptance |
|---|---|---|---|---|---|
| Signal-NVFP4, low | 100 % (30/30) | 302 | 107.1 | — | 5.93 |
| Signal-NVFP4, medium | 100 % (50/50) | 363 | 157.6 | 2.30 | 5.78 |
| GAIN, medium | 96 % (48/50) | 332 | 138.0 | 2.41 | 5.06 |
Read it as: at the same effort, Signal gives +14.2% decode via +14.2% acceptance and
−4.6% time per problem, with equal-or-better accuracy. And reasoning_effort=low
(which carries the only brevity instruction in the stock Qwen3.8 template; medium
is a no-op there) cuts Signal to 302 tokens/problem at 100% held.
Agentic code task (2048 engine, 3 files, hidden acceptance tests, temp 0.7)
| Success | Turns | Tokens | Wall | DFlash2 acceptance | |
|---|---|---|---|---|---|
| Signal-NVFP4 | 2/2 | 2 | 6 803 | 41.8 s | 6.33 (+4.2%) |
| GAIN | 2/2 | 2 | 6 950 | 42.7 s | 6.07 |
Both models solve the task in the same number of turns. Signal is marginally faster with a measurably better-aligned drafter.
Open conversation (15 general prompts, temp 0.7, deterministic)
| GAIN | Signal-NVFP4 | |
|---|---|---|
| total tokens | 8 985 (FR) / 7 671 (EN) | 12 235 / 10 534 (+36–37 %) |
Honest caveat: on open chat, this build is more verbose than the GAIN package. Upstream reports −57% answer tokens against the stock model — against an already-terse finetune, the advantage disappears on conversational prompts and shows up only on structured/reasoning work. Use GAIN for chat, Signal for code and reasoning.
Run (llama.cpp CUDA, RTX 5090)
llama-server.exe \
-m ...-NVFP4.gguf \
-md ...-draft-DFlash2-Q4_K_M.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 \
-ngl all -np 1 -c 262144 -fa on \
-ctk q8_0 -ctv q8_0 \
--spec-draft-type-k f16 --spec-draft-type-v f16 \
--chat-template-kwargs "{\"reasoning_effort\":\"low\"}" \
--cache-ram 32768 --no-cache-idle-slots \
--alias signal38-ultra:latest --reasoning-format auto
# add --mmproj ...-mmproj-BF16.gguf for image input (disables prompt-cache reuse)
Sampling — read this first. Use temperature 0.6 coding / 0.7 general, top_p 0.95, top_k 20. Do not use greedy (temperature 0): upstream documents degeneration on long generations at greedy, and one community report describes a repeat loop after 15–30 minutes on a from-scratch game build. For long builds also give the agent 64k+ context — a 30-minute build overflows 32k, and once the window shifts the model loses the start of its own code.
Reasoning effort: low is the recommended default here (fewest tokens at equal
accuracy in our GSM8K sweeps). medium decodes slightly faster per token on math.
xhigh is the template default; it was not re-measured for this build.
VRAM (measured, NVFP4 + DFlash2)
| Config | Slots | Context per slot | VRAM |
|---|---|---|---|
| Single agent, max context | 1 | 262 144 | 29 447 MiB |
Two-slot and f16-KV variants follow the GAIN package's table; the envelope is within 0.5 GB of it (identical file size).
Built from
- Base:
Qwen/Qwen3.8-27B - Finetune:
agentionai/Signal-3.8-27B("directness" self-distill; onlylm_head.weightdiffers from stock) - NVFP4 quant: converted from BF16 with
convert_hf_to_gguf.py@f3f1a8f27, thenllama-quantizewith a per-tensor map over a 200-chunk imatrix (NVFP4 backbone, Q8_0 heads). The recipes that putlm_headortoken_embdat 6-bit or below visibly damage this model — both were measured and rejected. - Drafter:
z-lab/Qwen3.8-27B-DFlash2(block-diffusion, lossless). Speculative decoding is lossless (rejection sampling), so the drafter never changes outputs, only speed.
Notes
- Thinking model — set
max_tokens ≥ 4096on hard tasks. - 262K context works; vision pairs with the included
mmproj-BF16.gguf. - CUDA graphs stay ON (measured +33% decode on this box in the sibling package).
- Ollama: see the GAIN sibling repo's
Modelfilefor the template (same one works here). - Only interleaved A/B runs are trusted here; see the sibling repo for the methodology
and the harnesses (
ab-run-07.ps1,eval-gsm8k-07.py,agentic.py).