neroued/Qwen3.8-27B-nvfp4-NInfer

🤗 Hugging Face 来源image-text-to-textapache-2.0激活 27B93 GBother✓ 1 个校验和今天更新
需要做种者 →

Qwen3.8-27B NVFP4 for NInfer

This model card is the version-controlled source for neroued/Qwen3.8-27B-nvfp4-NInfer.

The repository contains a mixed NVFP4/FP8 representation of Qwen3.8-27B. It combines the official BF16 checkpoint with the fixed packed Text weights from unsloth/Qwen3.8-27B-NVFP4 in the native NInfer .ninfer artifact format. The artifact is intended only for NInfer; it is not a Transformers checkpoint, Safetensors distribution, or GGUF file.

The artifact uses the Qwen3.5 Dense architecture. Text layers 0–55 use NVFP4 MLP weights, while the token embedding, attention input/output projections, GDN Q/K/V/Z and output projections, full output head, and Text layers 56–63 MLP weights use row-scaled FP8. Control weights use BF16, with separate MTP, Vision and DFlash2 weights.

Artifact

Field Value
Filename qwen3_8_27b_nvfp4.ninfer
Size 23,719,715,844 bytes (22.09 GiB)
SHA-256 74d2c57145e6ff11d1d2faa79594477f9bc903a611af1fb20218189fbbb77d82
Container version 3
Architecture Qwen3_5ForCausalLM
Public model name qwen3.8-27b
Chat template qwen3_8.jinja; override with --chat-template FILE
Template defaults thinking on; effort xhigh; closed-turn reasoning retained
Stored objects 1,246 (1,240 tensors and 6 resources)
NVFP4 tensors 112
Row-scaled FP8 tensors 146

The file contains Text, Vision, MTP, DFlash2, the optimized proposal head and frontend resources. Vision and speculative weights are loaded only when selected at startup. Source-derived NVFP4 and FP8 words are preserved without decode and requantization; only the official BF16 token embedding is encoded locally as row-scaled FP8.

Verify a downloaded file with:

printf '%s  %s\n' \
  '74d2c57145e6ff11d1d2faa79594477f9bc903a611af1fb20218189fbbb77d82' \
  'qwen3_8_27b_nvfp4.ninfer' | sha256sum --check

This release includes the complete DFlash2 companion weights from z-lab/Qwen3.8-27B-DFlash2 at revision 50307d4c4cde6860d4eee73e2547cd786fe8e8a4. Select --spec dflash2 --draft-tokens 7 --lm-head-draft; draft counts 1..15 are supported. DFlash2 requires the runtime revision listed below. Existing performance and evaluation tables retain their stated MTP configurations and revisions.

Requirements

  • NInfer revision 04350ba9 or later, built from source;
  • 64-bit Linux;
  • NVIDIA GeForce RTX 5090 (sm_120a);
  • CUDA Toolkit 13.1 or newer.

Already have the official v2 file? Upgrade it locally without downloading the weights again.

NInfer does not provide an install target or packaged binary. See the repository README for source-build dependencies.

Download and run a CLI example

hf download neroued/Qwen3.8-27B-nvfp4-NInfer \
  qwen3_8_27b_nvfp4.ninfer \
  --local-dir models

./build/apps/ninfer models/qwen3_8_27b_nvfp4.ninfer \
  --prompt "Explain prefill and decode in three sentences." \
  --max-context 32768 \
  --max-new 8192 \
  --kv-dtype fp8 \
  --spec mtp --draft-tokens 3 \
  --lm-head-draft

For images, videos, and structured chat history, see the CLI guide.

Start a local server

./build/apps/ninfer-serve models/qwen3_8_27b_nvfp4.ninfer \
  --host 127.0.0.1 \
  --port 8080 \
  --max-context 240000 \
  --kv-capacity 240000 \
  --max-concurrency 2 \
  --kv-dtype fp8 \
  --device-state-slots 2 \
  --host-state-slots 8 \
  --host-kv-mib 8192 \
  --spec mtp --draft-tokens 3 \
  --lm-head-draft \
  --preserve-thinking

Each request has a 240,000-token logical ceiling. The shared 240,000-token Device KV pool admits two active requests when their combined completion reservations fit; either request may use the full pool while running alone. Two extra Device checkpoint slots, eight pinned Host State slots, and 8 GiB of pinned Host KV retain reusable continuations under resource pressure.

See the HTTP serving guide for the API surface and the resource scheduling reference for cache and admission semantics.

Supported use

The artifact supports:

  • text generation in thinking and non-thinking modes;
  • image, multi-image, video, and mixed multimodal messages;
  • MTP speculative decoding with draft windows from one to five;
  • DFlash2 with draft windows from one to fifteen using the included DFlash2 companion weights (--spec dflash2 --draft-tokens 7, optionally --lm-head-draft);
  • BF16, INT8, FP8, NVFP4, and K8V4 KV cache;
  • CUDA Graph decode and compatible-prefix reuse;
  • startup-bounded small-scale concurrent serving with true batched decode;
  • the NInfer CLI;
  • OpenAI Responses Core, OpenAI Chat Completions, and Anthropic Messages serving.

Performance

Measured on September 28–29, 2026 with NInfer revision 7f6aafed, one RTX 5090, driver 617.14, and CUDA 13.4 compile/runtime/driver API. These serving runs use FP8 E4M3 row-256 KV, CUDA Graphs, a 1,024-token prefill chunk, disabled prefix reuse, and temperature 0.6 / top-p 0.95 / top-k 20 / min-p 0 / presence penalty 1.0 / frequency penalty 0. MTP0 has a 262,144-token context ceiling; MTP3 uses 131,072 tokens per request, three draft tokens, the optimized proposal head, and automatic shared KV capacity.

Concurrent MTP=3 corpus makespan

Each C is one complete 75-request corpus, with three reasoning and twelve cross-scenario fixtures, five seeds per fixture, and a fixed shuffled send order. Makespan includes prefill, decode, admission waits, transitions, and drain; actual output lengths vary.

C Requests Computed prefill tokens Decode tokens Makespan (s) Requests/s Corpus prefill (tok/s) Corpus decode (tok/s) Avg batch MTP acceptance
1 75 15,460 693,701 4,115.22 0.0182 3.8 168.6 1.00 59.6%
2 75 15,460 709,989 2,394.23 0.0313 6.5 296.5 1.90 58.9%
4 75 15,460 722,202 1,640.17 0.0457 9.4 440.3 3.13 58.5%
8 75 15,460 708,589 1,443.37 0.0520 10.7 490.9 3.52 59.5%

All 300 requests completed without request, CUDA, or allocation errors. At C=1/2/4/8, automatic KV capacity is 131,072 / 262,144 / 253,632 / 225,024 tokens. C=8 has average batch 3.52 and up to five waiting requests in the sampled intervals. The resident MTP3 weights occupy 19.729 GiB, and the workspace arena is 243.3 MiB.

Long-context serving (MTP disabled)

Values are arithmetic mean ± sample standard deviation over five fixed seeds per fixture.

Prompt tokens Samples Prefill phase (tok/s) Server TTFT (ms) Decode phase (tok/s)
7,680 5 12,819.1 ± 16.8 602.7 ± 1.1 74.1 ± 0.3
64,512 5 8,658.2 ± 44.8 7,487.4 ± 37.7 68.2 ± 0.2
130,048 5 6,198.5 ± 19.6 21,055.7 ± 67.2 62.5 ± 0.3
260,096 5 4,016.4 ± 10.4 64,910.0 ± 171.8 53.4 ± 0.6

MTP=3 single-request long-reasoning decode

The C=1 corpus supplies these phase statistics, with five samples per reasoning fixture.

Fixture Samples Completion tokens Decode phase (tok/s) MTP3 acceptance MTP3 tokens/round
long_decode_aime26_01 5 1,559.4 ± 727.2 207.4 ± 3.6 77.2% ± 1.8% 3.32 ± 0.05
long_decode_aime26_15 5 65,536.0 ± 0.0 161.7 ± 3.8 57.2% ± 2.1% 2.72 ± 0.06
long_decode_aime26_30 5 37,978.0 ± 7,474.4 170.0 ± 2.3 59.9% ± 1.5% 2.80 ± 0.05

MTP=3 single-request cross-scenario decode

Each category pools three fixtures × five seeds. Values are mean ± sample standard deviation.

Category Samples Decode phase (tok/s) MTP3 acceptance MTP3 tokens/round
Code 15 205.3 ± 9.5 76.8% ± 5.1% 3.30 ± 0.15
Story 15 132.6 ± 13.1 37.6% ± 7.1% 2.13 ± 0.21
Translation 15 202.9 ± 12.3 75.2% ± 6.6% 3.26 ± 0.20
Structured 15 231.7 ± 10.6 90.6% ± 5.6% 3.72 ± 0.17

All five AIME 15 samples reach the 65,536-token output budget. The C=1 corpus contains 40 stop-token and 35 output-limit results; all are retained in the statistics. These measurements do not score answer accuracy or task completion.

The full results and reproduction commands also cover DFlash2 K=7, MTP3 decode saturation, and completion outcomes.

Evaluation

The artifact was evaluated through NInfer's OpenAI-compatible serving route with thinking enabled, MTP=3, and INT8 group-64 KV. EvalScope 1.9.0 used 0-shot prompts, rule-based scoring, and one sample per problem with temperature 1.0, top-p 0.95, top-k 20, presence penalty 0.0, and seed 42. The text suite ran at a 252,928-token context limit; the multimodal suite ran with --vision at a 81,920-token limit.

Benchmark NInfer NVFP4 Correct / total Official Qwen3.8-27B BF16
IFBench (prompt-level strict) 77.00% 231 / 300 79.5
AIME 2025 96.67% 29 / 30 —
AIME 2026 96.67% 29 / 30 —
GPQA-Diamond 90.40% 179 / 198 89.2
ERQA 66.25% 265 / 400 65.5
RealWorldQA 83.53% 639 / 765 85.9

All 1,723 configured samples completed and were scored. IFBench additionally reports 80.50% instruction-level strict, 80.33% prompt-level loose, and 83.50% instruction-level loose. These are single-sample results, not pass@k.

The official Qwen3.8-27B BF16 figures come from the upstream model card; its sampling settings and IFBench metric level are not stated there, so the last column is not a same-protocol comparison. The NVFP4 deltas stay within ±2.5 points on the four overlapping benchmarks, and the upstream card reports no AIME results.

Limits

  • NInfer executes on one RTX 5090 and one CUDA device, with a startup-fixed capacity of 1–8 active requests per Engine.
  • It does not provide large-scale or preemptive continuous batching, priority/QoS scheduling, multi-GPU execution, CPU/GPU offload, or distributed serving.
  • Context allocation is subject to GPU memory and the selected KV-cache type.
  • NInfer does not execute generated tool calls.

Provenance

Field Value
Base repository Qwen/Qwen3.8-27B
Base revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
Base download source modelscope.cn/models/Qwen/Qwen3.8-27B
Quantized source repository unsloth/Qwen3.8-27B-NVFP4
Quantized source revision 60e813d4dbbdc5d64cf3f5a8caf2897bedf03679
Conversion recipe qwen3_8_27b_nvfp4
Embedding encoder fp8_row_maxabs
Converter repository https://github.com/Neroued/ninfer
Minimum runtime revision 98dada0e03cb073fd07f905400b5904bc6e82759
Ranking input SHA-256 c692dc76388132c910547589b4fb4a0503fbd6ad50aaac6a509bbcb192a8afa5

The artifact identity, summarized object inventory, and conversion provenance are published in artifact-manifest.json. The exact storage contract is maintained in the v3 container reference.

License

This NInfer artifact is distributed under the Apache License 2.0. The Qwen3.8-27B base repository and the quantized source repository are also licensed under Apache-2.0. Users remain responsible for complying with the license and applicable laws.