nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-MTP-NVFP4

🤗 Hugging Face sourcetext-generationapache-2.0182 GBsafetensorsHF checksums availableupdated today
No torrent yet

Qwen3.8-27B EfficientThink — NVFP4 + BF16 MTP

BF16 / FP8 main model · GGUF variants

True static FP8 DFlash2 draft

The bundled DFlash2 draft is now a pre-quantized static FP8 compressed-tensors checkpoint, not a BF16 checkpoint carrying an FP8 directory label. model.safetensors is 2,407,027,720 bytes (SHA256 1f3636a32d866f8ebc7f422d63f9247126ebb6d2566d3e0da327d81dd8fa25d1). Its audited tensor set contains 20 FP8 E4M3 weights with 20 FP32 scales and 61 retained BF16 tensors.

Load it explicitly with:

--speculative-draft-model-quantization compressed-tensors

Matched DGX Spark checks used the same W4A4 target, 15 prompts, XH, 256 generated tokens, and 8 draft tokens. All 15 requests completed in every cell.

Draft Concurrency Aggregate tok/s DFlash acceptance Mean accepted / 8 Request errors
BF16 reference C1 29.75 38.77% 3.72 0
Static FP8 C1 30.26 35.86% 3.51 0
BF16 reference C4 68.50 32.71% 3.29 0
Static FP8 C4 75.94 33.55% 3.35 0

C1 acceptance is the mean of same-run log snapshots; C4 acceptance is the post-run SGLang metrics gauge. The C4 static draft improved aggregate throughput by about 10.9% over BF16 in this matched check. This short fixed-length test validates serving behavior; it does not replace the formal capability scores elsewhere in this card.

Research disclaimer: This experimental release is provided solely to study the technical feasibility and behavioral effects of refusal-tendency dissolution. It is not a comprehensive safety conclusion, an endorsement of unrestricted use, or professional advice. Users are responsible for lawful and appropriate use and for independently verifying model outputs.

The figure presents this model's W4A4, W4A4+W8A8, W4A16, and W8A16 variants. Fast/Mixed use C24 and W4A16 uses C16; W8A16 uses C16 for GPQA/LCB and C24 for MMLU. All 12 / 11 GPQA timeouts remain in the Fast/Mixed 198-question failure denominators; no anomalous question was dropped.

True-QAT INT8 W8A8 | dynamic INT8 activations

Download directory: INT8-W8A8-QAT/. Matching main/standalone repository on this platform.

Files and precision

Component Path Precision / role Size
Main model INT8-W8A8-QAT/model-00001-of-00008.safetensorsmodel-00008-of-00008.safetensors True-QAT INT8 W8A8; dynamic INT8 activations 29.48 GB
Vision + native MTP INT8-W8A8-QAT/vision-mtp-bf16.safetensors 333 BF16 vision tensors + 15 BF16 MTP tensors, with 348 real index mappings 1.77 GB
Complete inventory INT8-W8A8-QAT/manifest.json and INT8-W8A8-QAT/SHA256SUMS Roles, bytes, and SHA256 for the current 29-file directory
Structured evaluation INT8-W8A8-QAT/evaluation/formal-quality-and-performance.json Formal scores, reasoning statistics, and the complete research record

Training and export method

  • 64-layer Qwen3.8-27B multimodal architecture with 1,599 entries in the published model index.
  • 3,200 QAT optimizer steps; all 400 language linear tensors recorded non-zero gradients and are published with INT8 weights.
  • Dynamic INT8 activations; 247 items entered the accepted training set.
  • BF16 scales were losslessly exported as F32; the vision tower and native MTP remain BF16.

Formal capability and reasoning results

Protocol: 1× RTX PRO 6000 Blackwell 96GB, vLLM 0.28.0 + native MTP3, C20, BF16 KV, reasoning_effort=xhigh, a 32,768-token output cap, and a 1,800-second request timeout. The long-output formal suite uses C20 because C24 did not leave enough KV capacity for the full suite.

Suite Score Mean reasoning P50 / P90 >8K / >16K 32K trunc. Empty final / unparseable
GPQA 162/198 (81.82%) 9,530 4,699.5 / 32,767 73 / 41 23 23 / 23
MMLU 451/500 (90.20%) 837 203 / 1,906.1 9 / 3 0 0 / 0
LCB 73/100 (73.00%) 13,921 8,347 / 32,768 50 / 41 23 23 / 23

Request / HTTP / capture / grader errors are all 0. LCB has 0 code timeouts and 0 syntax errors, plus 1 runtime error. IPC-v4 uniformly regraded the original 100 answers without issuing new model requests.

24 short-output cells for bundled runtime paths

Protocol: 1,024 input / 256 output, warmup plus 3 trials. This measures short fixed-length serving throughput, not long-reasoning speed. All 24 bare/MTP3 cells completed with 0 request errors.

  • Highest measured throughput for this tier: vLLM MTP3 C24 at 662 tok/s, 56.28% acceptance, about 27.6 tok/s/request.
  • SGLang MTP3 C24: 654 tok/s at 55.17% acceptance.
  • The release bundles and recommends native MTP3 only; the structured evaluation file preserves the complete historical research record.
Framework / mode C Aggregate tok/s Per-request tok/s Acceptance TTFT P50 Latency P50 Peak GPU Errors
vLLM bare C1 32 31.9 0.16s 8.03s 85.8 GiB 0
vLLM bare C4 113 28.2 0.56s 9.04s 86.1 GiB 0
vLLM bare C8 213 26.6 1.01s 9.55s 86.1 GiB 0
vLLM bare C16 363 22.7 1.59s 11.15s 86.1 GiB 0
vLLM bare C20 423 21.2 1.87s 11.95s 86.1 GiB 0
vLLM bare C24 476 19.8 2.16s 12.71s 86.1 GiB 0
vLLM MTP3 C1 56 55.9 47.48% 0.18s 4.58s 85.8 GiB 0
vLLM MTP3 C4 202 50.4 54.87% 0.56s 4.62s 86.0 GiB 0
vLLM MTP3 C8 356 44.5 57.08% 1.09s 5.27s 86.0 GiB 0
vLLM MTP3 C16 544 34.0 57.31% 1.70s 6.88s 86.0 GiB 0
vLLM MTP3 C20 619 31.0 56.30% 2.00s 7.77s 86.0 GiB 0
vLLM MTP3 C24 662 27.6 56.28% 2.31s 8.63s 86.0 GiB 0
SGLang bare C1 45 45.2 0.14s 5.67s 87.9 GiB 0
SGLang bare C4 158 39.4 0.43s 6.49s 88.1 GiB 0
SGLang bare C8 287 35.9 0.69s 7.13s 88.1 GiB 0
SGLang bare C16 469 29.3 1.22s 8.73s 88.1 GiB 0
SGLang bare C20 537 26.8 1.48s 9.53s 88.1 GiB 0
SGLang bare C24 597 24.9 1.74s 10.29s 88.1 GiB 0
SGLang MTP3 C1 86 86.1 67.86% 0.15s 2.97s 86.5 GiB 0
SGLang MTP3 C4 234 58.5 54.03% 0.43s 3.83s 86.7 GiB 0
SGLang MTP3 C8 379 47.4 51.95% 0.72s 4.96s 86.7 GiB 0
SGLang MTP3 C16 578 36.1 55.06% 1.25s 6.56s 86.7 GiB 0
SGLang MTP3 C20 612 30.6 54.42% 1.52s 8.06s 86.7 GiB 0
SGLang MTP3 C24 654 27.3 55.17% 1.79s 8.90s 86.7 GiB 0

Verified launch paths

cd INT8-W8A8-QAT
bash scripts/serve-vllm-mtp3.sh
Script Purpose
scripts/serve-vllm-bare.sh vLLM bare
scripts/serve-vllm-mtp3.sh vLLM native MTP3; recommended throughput path
scripts/serve-sglang-bare.sh SGLang bare
scripts/serve-sglang-mtp3.sh SGLang native MTP3

All four bundled paths passed text, image, and real-video smoke on the same model hash. SGLang compressed-tensors INT8 on Blackwell SM120/121 uses the bundled runtime/sglang-sm120-int8-compat/ compatibility layer.

Formal capability results

Variant GPQA 198 MMLU 500 LCB 100
W4A4 (NVFP4 Fast) 158/198 (79.80%) 447/500 (89.40%) 74/100 (74.00%)
W4A4+W8A8 (NVFP4 Mixed Precision) 168/198 (84.85%) 458/500 (91.60%) 75/100 (75.00%)
W4A16 161/198 (81.31%) 457/500 (91.40%) 74/100 (74.00%)
W8A16 159/198 (80.30%) 450/500 (90.00%) 78/100 (78.00%)

Fast/Mixed use the same C24 runtime protocol for all three scores. The 12 / 11 GPQA timeouts remain in the complete denominators; anomalous questions were neither removed nor rescored. Results from different runtime formats and decoders are not controlled measurements of quantization loss.

Mixed precision versus fast: scores and reasoning cost

Metric W4A4 Fast W4A4 + W8A8 Mixed Change
GPQA score 158/198 (79.80%) 168/198 (84.85%) +10 correct / +5.05pp
GPQA mean reasoning 9,123 8,400 -7.9%
GPQA P50 / P90 4,928.5 / 24,893.0 4,278.0 / 24,402.6
MMLU score 447/500 (89.40%) 458/500 (91.60%) +11 correct / +2.20pp
MMLU mean reasoning 963 848 -11.9%
MMLU P50 / P90 225.0 / 2,285.5 216.5 / 1,668.9
LCB score 74/100 (74.00%) 75/100 (75.00%) +1 correct / +1.00pp
LCB mean reasoning 14,602 13,647 -6.5%
LCB P50 / P90 9,626.5 / 32,769.0 7,441.5 / 32,769.0

Under this C24 protocol, mixed precision answers 10 more GPQA and 11 more MMLU questions correctly, while LCB accuracy is unchanged; mean reasoning tokens are lower in all three suites. This compares quantized variants, not the official base against post-training. MMLU/LCB means cover 500/100 questions; percentage changes use unrounded means.

W4A16 formal C16 full-suite results

Protocol: 1× RTX PRO 6000 96GB, SGLang + official BF16 MTP, C16, xhigh, a 32,768-token output cap, and a 1,800-second request timeout. This is not a controlled equal-concurrency comparison against the C24 W4A4 runs.

Suite Final score Mean reasoning P50 / P90 >8K / >16K 32K trunc. Other anomalies
GPQA 161/198 (81.31%) 11,404 6,795.5 / 32,767 92 / 57 27 29 empty final-channel outputs; 27 unparseable responses; 0 request errors
MMLU 457/500 (91.40%) 802 225 / 1,554 9 / 1 0 0 request errors; 0 empty finals
LCB 74/100 (74.00%) 14,511 8,513.5 / 32,769 51 / 42 26 0 request errors/timeouts; 26 empty-code cases

GPQA scoring: Final 161/198 (81.31%) across all 198 questions, with 0 request errors and 0 timeouts. Parsing accepts only a non-empty final channel or a complete explicit Final Answer: A/B/C/D on the last non-empty reasoning line.

LCB difficulty: Easy 23/23 (100%), Medium 29/31 (93.55%), Hard 22/46 (47.83%). All 26 length-limited outputs and all 26 empty-code cases remain failures in the 100-question denominator.

All three NVFP4 MMLU results use the same offline final-content-strict-single-letter-v2 rescore: Fast 447/500 (89.40%), Mixed Precision 458/500 (91.60%), and W4A16 457/500 (91.40%); generation outputs are unchanged. The retired 430/449/439 scores are not used.

W4A16 MTP short-output concurrency

The fixed short-output grid tested only C1/C4/C8/C16/C24; C2 and C32 were not tested. Aggregate throughput is rounded to whole tokens/s.

Concurrency Aggregate tok/s MTP acceptance Accepted draft / verification Committed output / verification
C1 127 72.84% 2.185 3.160
C4 452 70.20% 2.106 3.103
C8 764 71.04% 2.131 3.127
C16 (balanced) 1,216 74.28% 2.229 3.228
C24 (max throughput) 1,304 70.60% 2.118 3.114

All five cells had 0 request errors, 0 timeouts, and 0 empty outputs. Punctuation-collapse manual review was not part of this short sweep, so no zero claim is made for that field. C16 provides the highest acceptance while retaining 1,216 tok/s; C24 is the highest measured aggregate-throughput point.

The tested W4A16 SGLang settings are --quantization modelopt_mixed, EAGLE, steps=3, top-k=1, draft tokens=4, BF16 dtype/KV, and a 65,536-token context. The four-mode vLLM smoke still covers only Fast and Mixed Precision.

W8A16 quality-oriented FP8 package

The complete W8A16 package is available at W8A16/: 24 files / 38,477,562,434 bytes, including the manifests. Download the entire directory and do not mix it with W4A4/, W4A4+W8A8/, or W4A16/.

Format note: W8A16 uses block-wise FP8 E4M3 weights with BF16 activations and KV cache. It is grouped in this repository family for distribution, but it is not NVFP4 encoding.

  • 64-layer text trunk; 1,391 tensors in the complete package.
  • 192 MLP linear weights use FP8 E4M3 with 128×128 blocks; 305 other text linear weights remain BF16.
  • vision-mtp-bf16.safetensors is a 1,770,897,648-byte shared component containing the official 333 BF16 vision tensors and 15 BF16 MTP tensors. It is not a standalone main model.
  • The official MTP component was restored from Qwen revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0; it was not trained during this SFT/SimPO run.

W8A16 formal capability results

Protocol: one RTX PRO 6000 96GB, SGLang + official BF16 MTP, xhigh, 32,768-token output limit, and a 1,800-second request timeout. GPQA and LCB used C16; MMLU is the adopted clean C24 result.

Suite Final score Concurrency Request errors / timeouts
GPQA 159/198 (80.30%) C16 0 / 0
MMLU 450/500 (90.00%) C24 0 / 0
LCB 78/100 (78.00%) C16 0 / 0

For LCB, all 100 problems remain in the denominator: 22 length stops and the corresponding 22 empty-code submissions count as failures. The run recorded 0 request errors, 0 HTTP timeouts, 0 code-execution timeouts, and 0 syntax errors.

W8A16 MTP short-output concurrency

Measured on one RTX PRO 6000 96GB with SGLang + official MTP, xhigh, and one 256-token wave per cell. All 53/53 requests completed with 0 request errors.

Concurrency Aggregate tok/s MTP acceptance Accepted draft / verification Mean TTFT
C1 87 72.9% 2.19 0.063 s
C4 (highest acceptance) 326 75.8% 2.27 0.143 s
C8 562 75.7% 2.27 0.170 s
C16 955 74.7% 2.24 0.283 s
C24 (max throughput) 1,056 73.2% 2.20 0.257 s

This is a single-wave fixed-length sweep, not per-user speed or sustained 32K throughput. Spark SGLang/vLLM bare+MTP text/image/video smoke and PRO SGLang bare+MTP capability smoke also passed; these are short runtime compatibility checks, not formal general-quality scores. The adopted final GPQA/MMLU/LCB scores are listed above; no partial score is reported here.

Tested SGLang path

SGLANG_FORCE_FP8_MARLIN=1 python -m sglang.launch_server \
  --model-path ./W8A16 \
  --quantization modelopt_mixed \
  --dtype bfloat16 --kv-cache-dtype bfloat16 \
  --enable-linear-replayssm-spec \
  --speculative-algorithm EAGLE \
  --speculative-num-steps 3 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 4

The published metadata view above is the SGLang-tested path. vLLM bare+MTP smoke was validated only through a separate generic-FP8 metadata view of the same unchanged weights, with forced FP8 Marlin; that auxiliary view is not part of this download, so SGLang is the recommended published path.

MTP concurrency, speed and acceptance

This is the existing 256-token fixed-length short-output sweep: one wave per cell, 53 requests per variant. Aggregate throughput is neither per-user speed nor sustained 32K reasoning performance. The C16 cell is a completed short-output speed measurement, not a C16 full capability score.

Fast

Concurrency Aggregate tok/s MTP acceptance Accepted draft / verification Completion / verification
C1 82 75.64% 2.269 3.282
C4 318 76.56% 2.297 3.303
C8 688 74.49% 2.235 3.225
C16 1,196 74.25% 2.227 3.223
C24 1,684 73.17% 2.195 3.197

Mixed precision

Concurrency Aggregate tok/s MTP acceptance Accepted draft / verification Completion / verification
C1 102 72.92% 2.188 3.200
C4 374 73.08% 2.193 3.180
C8 665 72.21% 2.166 3.156
C16 1,176 72.71% 2.181 3.188
C24 1,595 73.42% 2.203 3.202

Both measured short-output grids reach their highest aggregate throughput at C24. C24 long-reasoning runs recorded timeouts, so C24 is not recommended as a validated long-output optimum. Acceptance is accepted / proposed draft tokens; it differs from accepted draft length and final completion length per verification. Dividing completion length by 4 does not give acceptance.

Choose and download

Directory Quantization Complete package size
W4A4/ W4A4 NVFP4, retained BF16 head/control components 20.62 GB
W4A4+W8A8/ W4A4 NVFP4 + W8A8 FP8, retained BF16 head/control components 25.80 GB
W4A16/ W4A16 ModelOpt NVFP4; BF16 activation/KV and official BF16 vision/MTP 20.62 GB

Choose mixed precision when prioritizing the capability scores in this evaluation. Fast is smaller and has higher throughput in this C24 short-output sweep. Package size is not VRAM usage. All four are safetensors packages, not GGUF; vision and native MTP remain BF16 rather than NVFP4. Each variant additionally bundles an optional true static FP8 DFlash2 draft under DFlash2-FP8/; DFlash2 and native MTP are alternative speculative decoders.

Download all 15 files in the chosen directory: model-nvfp4-fast.safetensors (W4A4/) or model-nvfp4-mixed.safetensors (W4A4+W8A8/), vision-mtp-bf16.safetensors, configs/index/tokenizer/processors, manifest.json, and SHA256SUMS. Do not download only the text shard or mix variant files. Each variant contains 64 text layers, 333 BF16 vision tensors, and 15 BF16 MTP tensors. The official MTP was not trained during this SFT/SimPO run.

AWQ-W4A16 | native MTP with vLLM

Complete directory: AWQ-W4A16/; the main weight is Qwen3.8-27B-EfficientThink-SimPO-AWQ-W4A16.safetensors. Download the entire directory; do not mix it with W4A4, W4A4+W8A8, the earlier W4A16 package, or W8A16.

Format identity: this is a compressed-tensors, pack-quantized AWQ W4A16 build. 367 target weights use asymmetric group-128 INT4, 33 target weights use symmetric group-128 INT8, and critical, vision, MTP, and other retained tensors remain BF16; activations and KV cache are BF16. It is not NVFP4 encoding, GPTQ, or imatrix. hf_quant_config.json is retained as upstream ModelOpt provenance; runtime loading follows config.json, where quant_method=compressed-tensors is authoritative.

vision-mtp-bf16.safetensors combines 333 official BF16 vision tensors and 15 official BF16 MTP tensors. It is the vision/MTP component for this model, not a DFlash2 draft.

Formal capability and reasoning results

Protocol: one RTX PRO 6000 96GB, vLLM + native MTP, C24, xhigh, a 32,768-token output cap, and a 1,800-second request timeout. Every anomalous sample remains in the denominator; anomaly categories may overlap.

Suite Final score Mean reasoning P50 / P90 >8K / >16K 32K trunc. Request errors / HTTP timeouts
GPQA 159/198 (80.30%) 11,142 6,283 / 32,768 87/198 (43.94%) / 56/198 (28.28%) 28/198 (14.14%) 0/198 / 0/198
MMLU 453/500 (90.60%) 935 218 / 1,674.4 12/500 (2.40%) / 4/500 (0.80%) 0/500 0/500 / 0/500
LCB 74/100 (74.00%) 13,922 8,207.5 / 32,768 50/100 (50.00%) / 42/100 (42.00%) 23/100 (23.00%) 0/100 / 0/100

GPQA has 28/198 (14.14%) empty-final, no-submission, and unparseable cases. MMLU has 0/500 empty, no-submission, and unparseable cases. For LCB, 22/100 (22.00%) no-code/no-submission cases, 23/100 (23.00%) unparseable outputs, and 1/100 (1.00%) syntax error remain failures; code-execution timeouts were 0/100. This vLLM C24 run is not a controlled quantization-loss comparison against the earlier SGLang W4A16, NVFP4, or other decoder results.

vLLM MTP short-output concurrency

The fixed workload used 1,024 input + 256 output tokens, three trials per cell, and num_speculative_tokens=3; all request-error counts were 0. TPS is rounded to whole tokens/s.

Concurrency Aggregate tok/s MTP acceptance
C1 39 56.03%
C4 128 51.82%
C8 226 52.19%
C16 366 49.05%
C24 (highest measured throughput) 480 53.46%

Verified launch path

python -m vllm.entrypoints.openai.api_server \
  --model ./AWQ-W4A16 \
  --served-model-name qwen38-27b-awq-w4a16 \
  --host 127.0.0.1 --port 19540 \
  --dtype bfloat16 \
  --quantization compressed-tensors \
  --kv-cache-dtype auto \
  --gpu-memory-utilization 0.95 \
  --max-model-len 65536 \
  --max-num-seqs 24 \
  --max-num-batched-tokens 2048 \
  --reasoning-parser qwen3 \
  --attention-backend TRITON_ATTN \
  --limit-mm-per-prompt '{"image":2,"video":1}' \
  --skip-mm-profiling \
  --no-enable-prefix-caching \
  --enforce-eager \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

vLLM 0.28.0 with compressed-tensors 0.17.0 and Transformers 5.15.1 passed bare/MTP text, Chinese, code, explanation, image, and real-MP4 smoke, including native MTP accepted/proposed counters. SGLang 0.5.19 with compressed-tensors 0.18.0, Transformers 5.12.1, FlashInfer 0.6.18, and Decord 0.6.0 also passed 6/6 in bare and MTP modes, but only inside an isolated overlay with a local CUDA libcudart link repair and an explicit Decord video-backend patch; this must not be read as stock-pip, drop-in compatibility. That SGLang run was a short smoke and provides no SGLang long-output quality or TPS claim. Reproducibility files are under AWQ-W4A16/runtime/.

AWQ-W4A16|vLLM 原生 MTP

完整目录:AWQ-W4A16/;主权重文件为 Qwen3.8-27B-EfficientThink-SimPO-AWQ-W4A16.safetensors。请下载整个目录,勿与 W4A4、W4A4+W8A8、旧 W4A16 或 W8A16 文件混用。

格式身份:这是 compressed-tensors 的 pack-quantized AWQ W4A16:367 个目标权重采用 group-128 非对称 INT4,33 个目标权重采用 group-128 对称 INT8,其余关键、视觉及 MTP 张量保留 BF16;激活与 KV cache 为 BF16。它不是 NVFP4 编码、GPTQ 或 imatrixhf_quant_config.json 仅保留上游 ModelOpt 来源记录,运行时以 config.jsonquant_method=compressed-tensors 为准。

vision-mtp-bf16.safetensors 同时包含 333 个官方 BF16 视觉张量和 15 个官方 BF16 MTP 张量;它是主模型的视觉/MTP 组件,不是 DFlash2 draft。

正式能力与思考量

口径:单张 RTX PRO 6000 96GB、vLLM + 原生 MTP、C24、xhigh、32,768 token 输出上限、1,800 秒请求超时。全部异常样本保留在分母;异常项可重叠。

项目 最终得分 平均思考 P50 / P90 >8K / >16K 32K 截断 请求错误 / HTTP 超时
GPQA 159/198(80.30%) 11,142 6,283 / 32,768 87/198(43.94%)/ 56/198(28.28%) 28/198(14.14%) 0/198 / 0/198
MMLU 453/500(90.60%) 935 218 / 1,674.4 12/500(2.40%)/ 4/500(0.80%) 0/500 0/500 / 0/500
LCB 74/100(74.00%) 13,922 8,207.5 / 32,768 50/100(50.00%)/ 42/100(42.00%) 23/100(23.00%) 0/100 / 0/100

GPQA 的空 final、无提交和不可解析均为 28/198(14.14%)。MMLU 的空答、无提交和不可解析均为 0/500。LCB 的无代码/无提交为 22/100(22.00%)、不可解析 23/100(23.00%)、语法错误 1/100(1.00%)、代码执行超时 0/100;均按失败计入。这里与旧 SGLang W4A16、NVFP4 或其他解码器结果不是同条件量化损失对照。

vLLM MTP 短输出并发

固定口径为 1,024 输入 + 256 输出,每档 3 次,num_speculative_tokens=3,所有请求错误为 0。TPS 按模型卡统一取整。

并发 聚合 tok/s MTP 接受率
C1 39 56.03%
C4 128 51.82%
C8 226 52.19%
C16 366 49.05%
C24(已测最高吞吐) 480 53.46%

已验证启动方式

python -m vllm.entrypoints.openai.api_server \
  --model ./AWQ-W4A16 \
  --served-model-name qwen38-27b-awq-w4a16 \
  --host 127.0.0.1 --port 19540 \
  --dtype bfloat16 \
  --quantization compressed-tensors \
  --kv-cache-dtype auto \
  --gpu-memory-utilization 0.95 \
  --max-model-len 65536 \
  --max-num-seqs 24 \
  --max-num-batched-tokens 2048 \
  --reasoning-parser qwen3 \
  --attention-backend TRITON_ATTN \
  --limit-mm-per-prompt '{"image":2,"video":1}' \
  --skip-mm-profiling \
  --no-enable-prefix-caching \
  --enforce-eager \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

vLLM 0.28.0 + compressed-tensors 0.17.0 + Transformers 5.15.1 已完成 bare/MTP 的文本、中文、代码、解释、图片和真实 MP4 smoke,并记录原生 MTP accepted/proposed 计数。SGLang 0.5.19 + compressed-tensors 0.18.0 + Transformers 5.12.1 + FlashInfer 0.6.18 + Decord 0.6.0 的 bare/MTP 同样各完成 6/6,但依赖隔离 overlay、局部 CUDA libcudart 链接修复与显式 Decord 视频后端补丁,不能理解为原生 pip 环境即装即用。SGLang 该轮只是短 smoke;没有 SGLang 长输出质量或 TPS 结论。可复现文件见 AWQ-W4A16/runtime/

Lynn Agent v0.86.7

Lynn Agent v0.86.7 now uses this release's Q2-LynnStyle / Q3-LynnStyle + DFlash2 packages. The pairing passed runtime validation on DGX Spark; macOS notarization, CI in both repositories, synchronization across all three release repositories, and public-network SHA verification also passed. Release-cache cleanup reclaimed approximately 4.01 GB, leaving approximately 22.34 GB free on the Tencent mirror disk; the current build, previous build, model files, and running service were not changed.

Installer China mirror GitHub fallback
Mac Apple Silicon Download Download
Mac Intel Download Download
Windows Download Download

Release records: primary GitHub repository · legacy GitHub repository · Gitee · CLI package

Known boundaries

C24 observation Fast Mixed precision
GPQA request timeouts 12 11
GPQA output-limit returns / empty final among returned requests 14 / 14 13 / 13
MMLU request errors / empty finals 0 / 0 0 / 0
LCB request errors 0 0
LCB output-limit returns / empty code 25 / 26 25 / 25

These counters overlap and must not be added as disjoint failures. Scores follow the frozen scorer's final decisions; an empty visible final is a separate observation. We do not claim zero strict loops or losslessness in every mode. A fast-variant non-thinking code probe answered incorrectly in both bare and MTP modes; its XH counterpart passed. The scored SGLang C24 evidence still covers the MTP path only; the separate vLLM 0.28.0 bare/MTP four-mode smoke is documented below.

Training method

Official Qwen3.8-27B → capability-preserving SFT → terminal-behavior SimPO → per-tensor FP32 delta merge → BF16 → three NVFP4 exports.

  • SFT: 1,905 samples, 1 epoch, 239 optimizer steps, effective batch 8; LoRA r16 / alpha32 / dropout 0.05; LR 5e-6 and 12 warmup steps.
  • SimPO: 110 preference pairs, 73 unique prompts, 5 optimizer steps; beta 1.0, gamma 0.2, peak LR 5e-7; LoRA r16 / alpha32 / dropout 0; two-GPU FSDP full sharding.
  • Training hardware: 2× RTX PRO 6000 Blackwell Server Edition.

SGLang runtime notes

SGLang native MTP retains the formal C24 quality and throughput evidence; vLLM 0.28.0 now has a separate four-mode startup/load/generation smoke. The frozen Spark launch records use the settings below; do not interchange the quantization backends:

Setting Fast Mixed precision
--quantization modelopt_fp4 modelopt_mixed
--speculative-algorithm EAGLE EAGLE
steps / top-k / draft tokens 3 / 1 / 4 3 / 1 / 4
KV / context BF16 / 65,536 BF16 / 65,536

That frozen environment uses SGLang source revision 17313cf4b25d with runtime adaptations; this is not a clean-upstream-install compatibility guarantee. The formal scores and short sweeps above were measured on PRO 6000, not Spark. Do not reuse GGUF DFlash2 arguments or private container-image names.

This page retains the C24 capability report.

vLLM 0.28.0 · tested on DGX Spark

Fast and mixed precision both completed real bare and native-MTP load, health, and text/image/video generation checks. The environment was DGX Spark GB10 (SM121) with the official Linux/ARM64 vLLM 0.28.0 image. Fast used modelopt_fp4; mixed precision used modelopt_mixed. Both selected FlashInferCutlassNvFp4LinearKernel for NVFP4, while mixed FP8 layers selected FlashInferFP8ScaledMMLinearKernel; neither fell back to Marlin.

Variant Decode Load memory Load time 6-case content check Image / video MTP accepted / drafted tokens
Fast bare 18.77 GiB 145.10 s 5/6 pass / pass
Fast MTP 19.56 GiB 198.43 s 5/6 pass / pass 101/168 (60.1%)
Mixed precision bare 23.52 GiB 143.88 s 5/6 pass / pass
Mixed precision MTP 24.31 GiB 211.90 s 6/6 pass / pass 112/168 (66.7%)

All six requests in every mode returned HTTP 200 with non-empty output and no repetitive-punctuation collapse. One short code check returned 55 instead of the expected 30 in fast bare/MTP and mixed bare, so those modes are reported as 5/6; mixed-precision MTP was 6/6. This was a short serial non-thinking smoke, not a vLLM TPS benchmark, general quality proof, 64K long-context test, or concurrency stress test. The 65,536 context and max-num-seqs=4 values were startup settings.

The recommended vLLM starting point is mixed precision + native MTP:

VARIANT=quality MODE=mtp PORT=19120 bash runtime/START_VLLM028_NVFP4.sh

The launcher binds only to 127.0.0.1, pins method=mtp and num_speculative_tokens=3, and refuses to pull an image or overwrite an existing container. Preload the image first; the script verifies the pinned official immutable manifest and local image ID. Model, quantization, and MTP arguments match the smoke above. The loopback host-network mapping is a deployment adapter for host access, not a new performance run.

Official references: vLLM ModelOpt quantization, vLLM MTP, and Docker host networking.


Qwen3.8-27B EfficientThink — NVFP4 + BF16 MTP

BF16 / FP8 主仓 · GGUF 版本

名副其实的静态 FP8 DFlash2 draft

仓内 DFlash2 draft 现已替换为预量化静态 FP8 compressed-tensors checkpoint,不再是放在 FP8 目录名下的 BF16 文件。model.safetensors2,407,027,720 bytes(SHA256 1f3636a32d866f8ebc7f422d63f9247126ebb6d2566d3e0da327d81dd8fa25d1);tensor 审计为 20 个 FP8 E4M3 权重、20 个 FP32 scale,以及 61 个保留 BF16 tensor。

加载时必须显式加入:

--speculative-draft-model-quantization compressed-tensors

同条件 DGX Spark 对照使用同一 W4A4 主模型、15 条固定输入、XH、256 输出 token 与 8 draft tokens;四个 cell 均为 15/15 请求成功。

Draft 并发 聚合 tok/s DFlash 接受率 平均接受长度 / 8 请求错误
BF16 对照 C1 29.75 38.77% 3.72 0
静态 FP8 C1 30.26 35.86% 3.51 0
BF16 对照 C4 68.50 32.71% 3.29 0
静态 FP8 C4 75.94 33.55% 3.35 0

C1 接受度为同轮服务日志快照均值,C4 接受度为结束后 SGLang metrics 精确值。本次同口径 C4 中,静态 FP8 draft 的聚合吞吐比 BF16 高约 10.9%。该短定长测试只验证服务行为,不替代本卡其他位置的正式能力分数。

科研免责声明: 本实验版本仅用于研究拒答倾向消融的技术可行性及行为影响,不构成全面的安全结论、对无限制使用的认可或专业建议。用户应依法、恰当地使用,并独立核验模型输出。

图中并列展示本模型的 W4A4、W4A4+W8A8、W4A16 与 W8A16。Fast/Mixed 为 C24,W4A16 为 C16;W8A16 的 GPQA/LCB 为 C16、MMLU 为 C24。GPQA 的 12 / 11 个超时均保留在 Fast/Mixed 的 198 题失败分母中,未丢弃异常题。

真 QAT INT8 W8A8|动态 INT8 激活

下载目录:INT8-W8A8-QAT/ 同平台对应主仓/独立仓

文件与精度

组件 路径 精度 / 角色 大小
主模型 INT8-W8A8-QAT/model-00001-of-00008.safetensorsmodel-00008-of-00008.safetensors 真 QAT INT8 W8A8;动态 INT8 激活 29.48 GB
视觉 + 原生 MTP INT8-W8A8-QAT/vision-mtp-bf16.safetensors 333 个 BF16 视觉张量 + 15 个 BF16 MTP 张量,索引真实引用 348 项 1.77 GB
完整清单 INT8-W8A8-QAT/manifest.jsonINT8-W8A8-QAT/SHA256SUMS 当前目录 29 个文件的角色、bytes 与 SHA256
结构化评测 INT8-W8A8-QAT/evaluation/formal-quality-and-performance.json 正式成绩、思考量与完整研究记录

训练与导出方法

  • 64 层 Qwen3.8-27B 多模态架构,发布包索引共 1,599 个张量。
  • 经过 3,200 个 QAT optimizer steps;400 个语言线性层均记录到非零梯度,并以 INT8 权重发布。
  • 激活采用动态 INT8;247 条训练样本进入已接受训练集。
  • BF16 scale 无损导出为 F32;视觉塔与原生 MTP 保持 BF16。

正式能力与思考量

协议:单张 RTX PRO 6000 Blackwell 96GB、vLLM 0.28.0 + 原生 MTP3、C20、BF16 KV、reasoning_effort=xhigh、32,768 输出上限、1,800 秒请求超时。正式长输出采用 C20,因为 C24 无法为整套长输出保留足够 KV 容量。

项目 得分 平均思考 P50 / P90 >8K / >16K 32K 截断 空 final / 不可解析
GPQA 162/198(81.82%) 9,530 4,699.5 / 32,767 73 / 41 23 23 / 23
MMLU 451/500(90.20%) 837 203 / 1,906.1 9 / 3 0 0 / 0
LCB 73/100(73.00%) 13,921 8,347 / 32,768 50 / 41 23 23 / 23

请求 / HTTP / capture / grader 错误均为 0;LCB 代码超时与语法错误均为 0,另有 1 个 runtime error。LCB 采用 IPC-v4 对原始 100 份回答统一复判,没有发出新模型请求。

24 个随包运行路径的短输出性能单元

协议:1,024 输入 / 256 输出、预热后 3 轮。下表只代表短定长服务吞吐,不代表长思考速度;24 个裸跑/MTP3 单元均为 0 请求错误。

  • 本档最高实测吞吐:vLLM MTP3 C24,662 tok/s,接受率 56.28%,约 27.6 tok/s/请求。
  • SGLang MTP3 C24:654 tok/s,接受率 55.17%。
  • 发布包只附带并推荐原生 MTP3;完整历史研究数据保留在结构化评测文件中。
框架 / 模式 并发 聚合 tok/s 每请求 tok/s 接受率 TTFT P50 延迟 P50 峰值显存 错误
vLLM 裸跑 C1 32 31.9 0.16s 8.03s 85.8 GiB 0
vLLM 裸跑 C4 113 28.2 0.56s 9.04s 86.1 GiB 0
vLLM 裸跑 C8 213 26.6 1.01s 9.55s 86.1 GiB 0
vLLM 裸跑 C16 363 22.7 1.59s 11.15s 86.1 GiB 0
vLLM 裸跑 C20 423 21.2 1.87s 11.95s 86.1 GiB 0
vLLM 裸跑 C24 476 19.8 2.16s 12.71s 86.1 GiB 0
vLLM MTP3 C1 56 55.9 47.48% 0.18s 4.58s 85.8 GiB 0
vLLM MTP3 C4 202 50.4 54.87% 0.56s 4.62s 86.0 GiB 0
vLLM MTP3 C8 356 44.5 57.08% 1.09s 5.27s 86.0 GiB 0
vLLM MTP3 C16 544 34.0 57.31% 1.70s 6.88s 86.0 GiB 0
vLLM MTP3 C20 619 31.0 56.30% 2.00s 7.77s 86.0 GiB 0
vLLM MTP3 C24 662 27.6 56.28% 2.31s 8.63s 86.0 GiB 0
SGLang 裸跑 C1 45 45.2 0.14s 5.67s 87.9 GiB 0
SGLang 裸跑 C4 158 39.4 0.43s 6.49s 88.1 GiB 0
SGLang 裸跑 C8 287 35.9 0.69s 7.13s 88.1 GiB 0
SGLang 裸跑 C16 469 29.3 1.22s 8.73s 88.1 GiB 0
SGLang 裸跑 C20 537 26.8 1.48s 9.53s 88.1 GiB 0
SGLang 裸跑 C24 597 24.9 1.74s 10.29s 88.1 GiB 0
SGLang MTP3 C1 86 86.1 67.86% 0.15s 2.97s 86.5 GiB 0
SGLang MTP3 C4 234 58.5 54.03% 0.43s 3.83s 86.7 GiB 0
SGLang MTP3 C8 379 47.4 51.95% 0.72s 4.96s 86.7 GiB 0
SGLang MTP3 C16 578 36.1 55.06% 1.25s 6.56s 86.7 GiB 0
SGLang MTP3 C20 612 30.6 54.42% 1.52s 8.06s 86.7 GiB 0
SGLang MTP3 C24 654 27.3 55.17% 1.79s 8.90s 86.7 GiB 0

已验证启动方式

cd INT8-W8A8-QAT
bash scripts/serve-vllm-mtp3.sh
脚本 用途
scripts/serve-vllm-bare.sh vLLM 裸跑
scripts/serve-vllm-mtp3.sh vLLM 原生 MTP3;推荐吞吐路径
scripts/serve-sglang-bare.sh SGLang 裸跑
scripts/serve-sglang-mtp3.sh SGLang 原生 MTP3

四条随包路径均在同哈希模型上通过文本、图片与真实视频 smoke。Blackwell SM120/121 的 SGLang compressed-tensors INT8 路径使用随包 runtime/sglang-sm120-int8-compat/ 兼容层。

三项正式成绩

版本 GPQA 198 MMLU 500 LCB 100
W4A4(NVFP4 极速版) 158/198(79.80%) 447/500(89.40%) 74/100(74.00%)
W4A4+W8A8(NVFP4 混合精度版) 168/198(84.85%) 458/500(91.60%) 75/100(75.00%)
W4A16 161/198(81.31%) 457/500(91.40%) 74/100(74.00%)
W8A16 159/198(80.30%) 450/500(90.00%) 78/100(78.00%)

Fast/Mixed 的三项成绩均使用 C24 运行协议。GPQA 的 12 / 11 个超时保留在完整分母中;未剔除或重算异常题。不同运行格式及解码器的结果不能直接当作等条件量化损失。

混合精度版相对极速版:得分与思考量

指标 W4A4 极速版 W4A4 + W8A8 混合精度版 变化
GPQA 得分 158/198 (79.80%) 168/198 (84.85%) +10题 / +5.05pp
GPQA 平均思考量 9,123 8,400 -7.9%
GPQA P50 / P90 4,928.5 / 24,893.0 4,278.0 / 24,402.6
MMLU 得分 447/500 (89.40%) 458/500 (91.60%) +11题 / +2.20pp
MMLU 平均思考量 963 848 -11.9%
MMLU P50 / P90 225.0 / 2,285.5 216.5 / 1,668.9
LCB 得分 74/100 (74.00%) 75/100 (75.00%) +1题 / +1.00pp
LCB 平均思考量 14,602 13,647 -6.5%
LCB P50 / P90 9,626.5 / 32,769.0 7,441.5 / 32,769.0

同一轮 C24 协议下,混合精度版 GPQA 多答对 10 题、MMLU 多答对 11 题、LCB 多答对 1 题;三项平均思考 token 均减少。这里比较的是量化版本,不是原版与后训练版。MMLU/LCB 均值分母为 500/100,百分比变化从未四舍五入的均值计算。

W4A16 正式 C16 全量结果

协议:单张 RTX PRO 6000 96GB、SGLang + 官方 BF16 MTP、C16、xhigh、32,768 输出上限、1,800 秒请求超时。此处与 W4A4 两版的 C24 不是严格等并发对照。

项目 最终得分 平均思考 P50 / P90 >8K / >16K 32K 截断 其他异常
GPQA 161/198(81.31%) 11,404 6,795.5 / 32,767 92 / 57 27 final 通道空 29;不可解析 27;请求错误 0
MMLU 457/500(91.40%) 802 225 / 1,554 9 / 1 0 请求错误 0;空 final 0
LCB 74/100(74.00%) 14,511 8,513.5 / 32,769 51 / 42 26 请求错误/超时 0;空代码 26

**GPQA 计分:**全部 198 题的最终成绩为 161/198(81.31%),请求错误 0、超时 0。解析仅接受非空 final 通道,或 reasoning 最后一个非空行中的完整明确 Final Answer: A/B/C/D

LCB 难度:Easy 23/23(100%)、Medium 29/31(93.55%)、Hard 22/46(47.83%)。26 个长度结束与 26 个空代码均按失败保留在 100 题分母。

三套 NVFP4 的 MMLU 均采用同一 final-content-strict-single-letter-v2 离线重算:极速版 447/500(89.40%)、混合精度版 458/500(91.60%)、W4A16 457/500(91.40%);生成内容未改变。旧的 430/449/439 分数不再使用。

W4A16 MTP 短输出并发

固定短输出网格仅测试 C1/C4/C8/C16/C24;C2 与 C32 未测。聚合吞吐按模型卡规则取整。

并发 聚合 tok/s MTP 接受率 接受草稿 / 验证轮 实际提交 / 验证轮
C1 127 72.84% 2.185 3.160
C4 452 70.20% 2.106 3.103
C8 764 71.04% 2.131 3.127
C16(平衡推荐) 1,216 74.28% 2.229 3.228
C24(最高吞吐) 1,304 70.60% 2.118 3.114

五档均为 0 请求错误、0 超时、0 空输出;本短扫未进行标点坍塌人工终审,因此不作对应零值声明。C16 在保持 1,216 tok/s 时接受率最高,作为平衡档;C24 是已测最大聚合吞吐。

W4A16 的已测 SGLang 参数为 --quantization modelopt_mixed、EAGLE、steps=3、top-k=1、draft tokens=4、BF16 dtype/KV、65,536 context;vLLM 四路 smoke 仍只覆盖极速版与混合精度版。

W8A16 质量优先 FP8 包

完整 W8A16 包已发布在 W8A16/:含清单共 24 个文件 / 38,477,562,434 bytes。请下载整个目录,不要与 W4A4/W4A4+W8A8/W4A16/ 混用。

格式说明: W8A16 使用分块 FP8 E4M3 权重 + BF16 激活与 KV cache。它为了统一分发放在本仓系列中,但编码格式不是 NVFP4

  • 64 层文字主干;完整包共 1,391 个张量。
  • 192 个 MLP 线性权重使用 128×128 分块 FP8 E4M3;其余 305 个文字线性权重保留 BF16。
  • vision-mtp-bf16.safetensors 为 1,770,897,648-byte 共用组件,包含官方 333 个 BF16 视觉张量与 15 个 BF16 MTP 张量;它不是可独立运行的主模型。
  • 官方 MTP 来自 Qwen revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0,未参与本轮 SFT/SimPO 训练。

W8A16 正式能力成绩

口径:单张 RTX PRO 6000 96GB、SGLang + 官方 BF16 MTP、xhigh、32,768 token 输出上限、1,800 秒请求超时。GPQA 与 LCB 使用 C16;MMLU 采用已冻结的 clean C24 结果。

项目 最终得分 并发 请求错误 / 超时
GPQA 159/198(80.30%) C16 0 / 0
MMLU 450/500(90.00%) C24 0 / 0
LCB 78/100(78.00%) C16 0 / 0

LCB 保留全部 100 题作为分母:22 次长度结束及对应的 22 个空代码均按失败计入;请求错误、HTTP 超时、代码执行超时和语法错误均为 0。

W8A16 MTP 短输出并发

实测环境:单张 RTX PRO 6000 96GB、SGLang + 官方 MTP、xhigh、每档单波 256-token 定长输出。53/53 个请求全部完成,请求错误 0。

并发 聚合 tok/s MTP 接受率 接受草稿 / 验证轮 平均 TTFT
C1 87 72.9% 2.19 0.063 秒
C4(接受率最高) 326 75.8% 2.27 0.143 秒
C8 562 75.7% 2.27 0.170 秒
C16 955 74.7% 2.24 0.283 秒
C24(最高吞吐) 1,056 73.2% 2.20 0.257 秒

这是单波定长短测,不是单用户速度,也不是 32K 持续吞吐。Spark 上的 SGLang/vLLM bare+MTP 文本/图片/视频 smoke,以及 PRO 上的 SGLang bare+MTP 能力 smoke 也已通过;它们只证明短请求运行兼容性,不是正式通用质量分数。已采用的 GPQA/MMLU/LCB 最终成绩见上方;本段不披露 partial 分数。

已测 SGLang 路径

SGLANG_FORCE_FP8_MARLIN=1 python -m sglang.launch_server \
  --model-path ./W8A16 \
  --quantization modelopt_mixed \
  --dtype bfloat16 --kv-cache-dtype bfloat16 \
  --enable-linear-replayssm-spec \
  --speculative-algorithm EAGLE \
  --speculative-num-steps 3 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 4

以上已发布元数据是 SGLang 实测路径。vLLM bare+MTP smoke 仅通过同一权重的独立通用 FP8 元数据视图并强制 FP8 Marlin 完成;该辅助视图未随本目录发布,因此公开推荐路径仍为 SGLang。

MTP 并发速度与接受率

以下是既有256-token 定长短输出测试,每档单批,两版各 53 个请求。聚合吞吐不是单用户速度,也不是 32K 长推理持续吞吐。这里的 C16 是已完成的短输出并发测速,不代表 C16 全量能力评分。

极速版

并发 聚合 tok/s MTP 接受率 接受草稿 / 验证轮 实际输出 / 验证轮
C1 82 75.64% 2.269 3.282
C4 318 76.56% 2.297 3.303
C8 688 74.49% 2.235 3.225
C16 1,196 74.25% 2.227 3.223
C24 1,684 73.17% 2.195 3.197

混合精度版

并发 聚合 tok/s MTP 接受率 接受草稿 / 验证轮 实际输出 / 验证轮
C1 102 72.92% 2.188 3.200
C4 374 73.08% 2.193 3.180
C8 665 72.21% 2.166 3.156
C16 1,176 72.71% 2.181 3.188
C24 1,595 73.42% 2.203 3.202

两版已测短输出网格均在 C24 取得最高聚合吞吐,但 C24 长推理出现过超时,因此暂不推荐它作为长输出最佳并发。接受率按实际接受草稿数 / 提议草稿数计算;它与每轮接受长度、每轮最终输出不同,不能用输出长度除以 4 替代。

选择与下载

目录 量化方案 完整包大小
W4A4/ W4A4 NVFP4,保留 BF16 头部与控制部件 20.62 GB
W4A4+W8A8/ W4A4 NVFP4 + W8A8 FP8,保留 BF16 头部与控制部件 25.80 GB
W4A16/ W4A16 ModelOpt NVFP4;BF16 activation/KV,官方 BF16 视觉与 MTP 20.62 GB

混合精度版更适合优先考虑本轮能力成绩的使用者;极速版文件较小,在本次 C24 短输出下吞吐更高。文件大小不是显存需求。四版都是 safetensors 包,不是 GGUF;视觉与原生 MTP 均为 BF16,而非 NVFP4。每个档位另在 DFlash2-FP8/ 提供可选的真静态 FP8 DFlash2 draft;DFlash2 与原生 MTP 是两条可选的投机解码路径。

下载所选目录的完整 15 个文件:model-nvfp4-fast.safetensorsW4A4/)或 model-nvfp4-mixed.safetensorsW4A4+W8A8/)、vision-mtp-bf16.safetensors、配置/索引/tokenizer/processor、manifest.jsonSHA256SUMS。不要只下载文字分片,也不要把三个目录的文件混放。每版含 64 层文字主干、333 个 BF16 视觉张量及 15 个 BF16 MTP 张量;官方 MTP 未参与本次 SFT/SimPO 训练。

Lynn Agent v0.86.7

Lynn Agent v0.86.7 已采用本系列 Q2-LynnStyle / Q3-LynnStyle + DFlash2。组合已在 DGX Spark 实测通过;Mac 公证、两仓 CI、三仓同步与公网 SHA 校验均通过。发布缓存清理后实际释放约 4.01 GB,腾讯镜像盘剩余约 22.34 GB;当前版、上一版、模型文件和服务均未改动。

安装包 国内镜像 GitHub 备用
Mac Apple Silicon 下载 下载
Mac Intel 下载 下载
Windows 下载 下载

发布记录:GitHub 主仓 · GitHub 旧仓 · Gitee · CLI 包

已知边界

C24 观测 极速版 混合精度版
GPQA 请求超时 12 11
GPQA 达到输出上限 / 已返回请求空 final 14 / 14 13 / 13
MMLU 请求错误 / 空 final 0 / 0 0 / 0
LCB 请求错误 0 0
LCB 达到输出上限 / 空代码 25 / 26 25 / 25

同一题可同时进入多个计数,不能相加为互斥失败数。得分采用冻结判分器最终结果,空可见 final 是单独的观测项。当前不宣称严格死循环为零或全模式无损。极速版不思考代码探针曾在 bare 和 MTP 模式均答错,XH 对应探针答对。SGLang C24 计分仍只对应 MTP 路径;另行完成的 vLLM 0.28.0 bare/MTP 四路 smoke 见下方。

训练方法

官方 Qwen3.8-27B → 能力保持 SFT → 终止行为 SimPO → 逐张量 FP32 增量合并 → BF16 → 三版 NVFP4 导出。

  • SFT:1,905 条样本,1 epoch,239 个优化器步,有效 batch 8;LoRA r16 / alpha32 / dropout 0.05;LR 5e-6、12 步 warmup。
  • SimPO:110 对偏好数据、73 个唯一 prompt,5 个优化器步;beta 1.0、gamma 0.2、峰值 LR 5e-7;LoRA r16 / alpha32 / dropout 0,双卡 FSDP full sharding。
  • 训练硬件:2× RTX PRO 6000 Blackwell Server Edition。

SGLang 运行说明

SGLang 原生 MTP 保留正式 C24 质量与吞吐数据;vLLM 0.28.0 已另行完成四路启动/加载/生成 smoke。冻结的 Spark 启动记录使用以下参数,不能把两版的量化后端混用:

设置 极速版 混合精度版
--quantization modelopt_fp4 modelopt_mixed
--speculative-algorithm EAGLE EAGLE
steps / top-k / draft tokens 3 / 1 / 4 3 / 1 / 4
KV / context BF16 / 65,536 BF16 / 65,536

该冻结环境使用 SGLang 源码版本 17313cf4b25d,并带运行环境适配;它不是干净上游安装的兼容性保证。上方正式得分和短扫则来自 PRO 6000,不能当作 Spark 的测速。不要照搬 GGUF DFlash2 参数或私有镜像名。

本页继续保留 C24 能力报告。

vLLM 0.28.0 · DGX Spark 实测

极速版与混合精度版均完成 bare 与原生 MTP 四路真实加载、健康检查和文本/图片/视频生成。实测环境为 DGX Spark GB10(SM121)、官方 Linux/ARM64 vLLM 0.28.0 镜像;极速版使用 modelopt_fp4,混合精度版使用 modelopt_mixed。三版 NVFP4 均实际选择 FlashInferCutlassNvFp4LinearKernel,混合精度版的 FP8 层使用 FlashInferFP8ScaledMMLinearKernel,未回退到 Marlin。

版本 解码 加载显存 加载时间 6 项内容校验 图片 / 视频 MTP 接受 token / 提议 token
极速版 bare 18.77 GiB 145.10 秒 5/6 通过 / 通过
极速版 MTP 19.56 GiB 198.43 秒 5/6 通过 / 通过 101/168(60.1%)
混合精度版 bare 23.52 GiB 143.88 秒 5/6 通过 / 通过
混合精度版 MTP 24.31 GiB 211.90 秒 6/6 通过 / 通过 112/168(66.7%)

四路各 6 个请求均为 HTTP 200、非空输出,未出现重复标点坍塌。极速 bare/MTP 与混合 bare 的同一道简短代码校验答为 55,正确值是 30,因此如实记为 5/6;混合精度 MTP 为 6/6。这里是关闭思考的短串行 smoke,不是 vLLM TPS、通用质量、64K 长上下文或并发压力证明;65,536 context 与 max-num-seqs=4 是启动配置值。

推荐的 vLLM 起步组合是混合精度版 + 原生 MTP

VARIANT=quality MODE=mtp PORT=19120 bash runtime/START_VLLM028_NVFP4.sh

启动器默认只监听 127.0.0.1,固定 method=mtpnum_speculative_tokens=3,并拒绝自动拉取镜像或覆盖同名容器。必须预先准备镜像;脚本会核对固定的官方不可变 manifest 与本地 image ID。模型/量化/MTP 参数来自上述实测;便于宿主访问的 loopback host-network 连接是部署适配,不作为新的性能实测。

官方参考:vLLM ModelOpt 量化vLLM MTPDocker host network