SergiioB/Qwen3.8-27B-int4-gdn8-ov

🤗 Hugging Face 来源image-text-to-textapache-2.0激活 27B26 GBother✓ 8 个校验和今天更新
需要做种者 →

Qwen3.8-27B OpenVINO INT4 (GDN / linear_attn kept out of INT4)

Unofficial export from Qwen/Qwen3.8-27B. Not a re-host of OpenVINO/Qwen3.8-27B-int4-ov.

Why this exists

The official Hub INT4 (OpenVINO/Qwen3.8-27B-int4-ov) is marked EXPERIMENTAL. Its openvino_config.json has:

  • dtype: int4_int4
  • quantization_configs.lm_model.ignored_scope: null
  • bits: 4, ratio: 1.0, sym: false

That puts GatedDeltaNet / linear_attn in INT4. Same class of bug as openvino.genai#3870 (optimum-intel dropped ignored scopes).

Measured on one Intel Arc Pro B70, OpenVINO GenAI 2026.5, official HF ai2d demo image, greedy:

Device Output
GPU.0 ")-+&-&)&/-,'/!&*%!'.33).)'03*#*,"
CPU (251 s) same string
Dummy zeros image all ! (token id 0)

Not a GPU plugin bug. CPU matches GPU. Fresh JIT cache, DYNAMIC_QUANTIZATION_GROUP_SIZE=0, official max_new_tokens= kwargs: still garbage.

Community copies (circulus/*-int4-ov, Morteza89/qwen3.8-27b-int4-ov) also ship ignored_scope: null. Do not download those as a fix.

Hub INT8 (OpenVINO/Qwen3.8-27B-int8-ov) is coherent on the same card.

What this repo is

Re-export from BF16:

  1. optimum / OpenVINO trace of Qwen/Qwen3.8-27B (Transformers 5.2.0 — 5.4.0 refuses the export)
  2. NNCF compress_weights on the language IR only:
    • INT4_ASYM, group 128, ratio 1.0
    • IgnoredScope(patterns=[".*linear_attn.*"])
  3. convert_tokenizer --with-detokenizer

NNCF bitwidth on the language graph:

Mode Share
INT4_ASYM g128 73% (256 / 545 layers)
float (ignored GDN) 22% (288 / 545)
INT8_ASYM per-channel 5% (1 layer)

Language bin 21 GB vs official all-INT4 14 GB. The extra 7 GB is GDN left at float.

Runtime

  • OpenVINO GenAI 2026.4+ (this IR is a VLM graph). Cascadia 0.2.3 stock ov-genai ships 2026.2 and SIGSEGVs in Tokenizer::setup_tokenizer.
  • Device: GPU.0, hint.performance_mode=LATENCY, hint.inference_precision=f16.
  • Call pipe.generate(prompt, image=..., generation_config=cfg) by keyword. A positional config is treated as images.
import numpy as np, openvino as ov, openvino_genai as og
from PIL import Image

P = ov.properties
pipe = og.VLMPipeline(
    "SergiioB/Qwen3.8-27B-int4-gdn8-ov",
    "GPU.0",
    **{
        P.hint.performance_mode: P.hint.PerformanceMode.LATENCY,
        P.hint.inference_precision: ov.Type.f16,
    },
)
cfg = og.GenerationConfig()
cfg.max_new_tokens = 128
image = ov.Tensor(np.array(Image.open("photo.jpg").convert("RGB"))[None])
print(pipe.generate("Describe this image.", image=image, generation_config=cfg))

B70 numbers (2026-09-12)

One Arc Pro B70, 150 W cap, GPU.0, greedy. Quality = English on the HF ai2d fire image.

Engine Artifact n=128 wall tok/s Quality
OpenVINO GenAI 2026.5 this INT4-GDN8 13.89 PASS
OpenVINO GenAI 2026.5 Hub INT8 13.09 PASS
OpenVINO GenAI 2026.5 Hub INT4 ~30 (token 0 / punctuation) FAIL
Cascadia 0.2.3 qwen35 Hub INT8 1-stage shard 4.56 PASS (host-side logits)
llama.cpp SYCL Q4_K_M GGUF 18.58 tg128 PASS (different artifact)

Do not put Hub INT4 tok/s on a leaderboard.

License

Apache-2.0, same as Qwen3.8-27B. Intel did not publish these files.