neroued/Qwen3.6-27B-NInfer

🤗 Hugging Face sourceimage-text-to-textapache-2.070 GBotherHF checksums availableupdated today
No torrent yet

Qwen3.6-27B for NInfer

This model card is the version-controlled source for neroued/Qwen3.6-27B-NInfer.

The repository contains Qwen3.6-27B converted to the native NInfer .ninfer artifact format. The artifact is intended only for NInfer; it is not a Transformers checkpoint, Safetensors distribution, or GGUF file.

Artifact

Field Value
Filename qwen3_6_27b.ninfer
Size 17,495,538,688 bytes (16.29 GiB)
SHA-256 9b610a7d051e7c4dbf89adb604bd269d248b643c8f6c7ef75bdb871af92c6f6b
Container version 3
Architecture Qwen3_5ForCausalLM
Public model name qwen3.6-27b
Chat template qwen3_6.jinja; override with --chat-template FILE
Template defaults thinking on; closed-turn reasoning omitted

The file contains Text, Vision, MTP, the optimized proposal head, and frontend resources. Text projections use Q4/Q5, vocabulary weights use Q6, and MTP projections use Q8. Vision and speculative weights are loaded only when selected at startup.

Verify a downloaded file with:

printf '%s  %s\n' \
  '9b610a7d051e7c4dbf89adb604bd269d248b643c8f6c7ef75bdb871af92c6f6b' \
  'qwen3_6_27b.ninfer' | sha256sum --check

Requirements

  • NInfer revision 98dada0 or later, built from source;
  • 64-bit Linux;
  • NVIDIA GeForce RTX 5090 (sm_120a);
  • CUDA Toolkit 13.1 or newer.

Already have the official v2 file? Upgrade it locally without downloading the weights again.

NInfer does not provide an install target or packaged binary. See the repository README for source-build dependencies.

Download and run a CLI example

hf download neroued/Qwen3.6-27B-NInfer \
  qwen3_6_27b.ninfer \
  --local-dir models

./build/apps/ninfer models/qwen3_6_27b.ninfer \
  --prompt "Explain prefill and decode in three sentences." \
  --max-context 32768 \
  --max-new 8192 \
  --kv-dtype fp8 \
  --spec mtp --draft-tokens 3 \
  --lm-head-draft

For images, videos, and structured chat history, see the CLI guide.

Start a local server

./build/apps/ninfer-serve models/qwen3_6_27b.ninfer \
  --host 127.0.0.1 \
  --port 8080 \
  --max-context 240000 \
  --kv-capacity 240000 \
  --max-concurrency 2 \
  --kv-dtype fp8 \
  --device-state-slots 2 \
  --host-state-slots 8 \
  --host-kv-mib 8192 \
  --spec mtp --draft-tokens 3 \
  --lm-head-draft \
  --preserve-thinking

Each request has a 240,000-token logical ceiling. The shared 240,000-token Device KV pool admits two active requests when their combined completion reservations fit; either request may use the full pool while running alone. Two extra Device checkpoint slots, eight pinned Host State slots, and 8 GiB of pinned Host KV retain reusable continuations under resource pressure.

See the HTTP serving guide for the API surface and the resource scheduling reference for cache and admission semantics.

Supported use

The artifact supports:

  • text generation in thinking and non-thinking modes;
  • image, multi-image, video, and mixed multimodal messages;
  • MTP speculative decoding with draft windows from one to five;
  • BF16, INT8, FP8, NVFP4, and K8V4 KV cache;
  • CUDA Graph decode and compatible-prefix reuse;
  • startup-bounded small-scale concurrent serving with true batched decode;
  • the NInfer CLI;
  • OpenAI Responses Core, OpenAI Chat Completions, and Anthropic Messages serving.

Performance

The single-request serving measurements below were collected on an NVIDIA GeForce RTX 5090 with CUDA 13.1. Requests were submitted serially to a persistent ninfer-serve process with CUDA Graph enabled, a 1,024-token prefill chunk, INT8 group-64 KV cache, and prefix reuse disabled. Each single-request value is the arithmetic mean ± sample standard deviation over five fixed seeds; warm-up requests are excluded. These results use the groupwise-int recipe.

Concurrent MTP=3 decode saturation

The concurrent campaign uses CUDA driver API 13.3 and one 293-token prompt followed by an 8,192-token generation per active request. Each concurrency point starts a fresh server with MTP3, INT8 group-64 KV, CUDA Graphs, a 16,384-token per-request context limit, and prefix reuse disabled. Aggregate throughput includes only complete one-second intervals whose actual decode batch remains equal to C. Each row is one sustained wave.

C Steady decode (tok/s) Speedup vs. C1 Wave makespan
1 185.8 1.00× 44.23 s
2 247.0 1.33× 66.67 s
4 309.5 1.67× 107.49 s
8 535.0 2.88× 125.20 s

Long-context baseline (MTP disabled)

Prompt tokens Prefill phase (tok/s) Server TTFT (ms) Decode phase (tok/s)
7,680 3,218.1 ± 4.3 2,392.4 ± 3.0 77.6 ± 0.1
64,512 2,655.9 ± 2.9 24,335.7 ± 25.2 70.7 ± 0.1
130,048 2,185.3 ± 0.3 59,590.3 ± 8.9 64.5 ± 0.1
260,096 1,614.8 ± 0.6 161,221.8 ± 62.5 54.8 ± 0.1

MTP=3 long-reasoning decode

Thinking was enabled and the output limit was 65,536 tokens.

AIME 2026 fixture Completion tokens Decode phase (tok/s) MTP acceptance MTP tokens/round
Problem 1 10,686.2 ± 553.8 175.4 ± 1.0 77.9% ± 0.9% 3.34 ± 0.03
Problem 15 61,604.2 ± 5,677.9 161.9 ± 2.8 73.4% ± 1.7% 3.20 ± 0.05
Problem 30 47,339.8 ± 9,162.2 172.2 ± 0.9 78.8% ± 0.8% 3.36 ± 0.02

MTP=3 cross-scenario decode

Each category contains three fixtures and five seeds per fixture (15 samples). Thinking was disabled and the output limit was 4,096 tokens.

Category Decode phase (tok/s) MTP acceptance MTP tokens/round
Code 167.0 ± 5.4 72.3% ± 3.5% 3.17 ± 0.11
Story 112.6 ± 9.4 37.8% ± 5.9% 2.13 ± 0.18
Translation 161.5 ± 11.3 68.3% ± 7.2% 3.05 ± 0.22
Structured output 193.0 ± 18.8 88.7% ± 11.7% 3.66 ± 0.35

See the full methodology and results, including metric definitions and the exact reproduction command.

Evaluation

The artifact was evaluated through NInfer's OpenAI-compatible serving route with thinking enabled, MTP=3, and a 262,144-token context limit. EvalScope 1.9.0 used 0-shot prompts, rule-based scoring, and one sample per problem with temperature 0.6, top-p 0.95, top-k 20, presence penalty 1.0, and seed 42. All 258 configured samples completed and were scored.

Benchmark Accuracy Correct / total
AIME 2025 86.67% 26 / 30
AIME 2026 93.33% 28 / 30
GPQA-Diamond 86.87% 172 / 198

These are single-sample results under the stated NInfer evaluation profile, not pass@k scores.

Limits

  • NInfer executes on one RTX 5090 and one CUDA device, with a startup-fixed capacity of 1–8 active requests per Engine.
  • It does not provide large-scale or preemptive continuous batching, priority/QoS scheduling, multi-GPU execution, CPU/GPU offload, or distributed serving.
  • Context allocation is subject to GPU memory and the selected KV-cache type.
  • NInfer does not execute generated tool calls.

Provenance

Field Value
Source repository Qwen/Qwen3.6-27B
Source revision 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Conversion recipe qwen3_6_27b
Converter repository https://github.com/Neroued/ninfer
Minimum runtime revision 98dada0e03cb073fd07f905400b5904bc6e82759

The artifact identity, summarized object inventory, and conversion provenance are published in artifact-manifest.json. The exact storage contract is maintained in the v3 container reference.

License

This NInfer artifact is distributed under the Apache License 2.0. The source Qwen3.6-27B repository is also licensed under Apache-2.0. Users remain responsible for complying with the license and applicable laws.