Qwen3.8-27B OpenVINO INT4 (GDN / linear_attn kept out of INT4)
Unofficial export from Qwen/Qwen3.8-27B. Not a re-host of OpenVINO/Qwen3.8-27B-int4-ov.
Why this exists
The official Hub INT4 (OpenVINO/Qwen3.8-27B-int4-ov) is marked EXPERIMENTAL. Its openvino_config.json has:
dtype: int4_int4quantization_configs.lm_model.ignored_scope: nullbits: 4,ratio: 1.0,sym: false
That puts GatedDeltaNet / linear_attn in INT4. Same class of bug as openvino.genai#3870 (optimum-intel dropped ignored scopes).
Measured on one Intel Arc Pro B70, OpenVINO GenAI 2026.5, official HF ai2d demo image, greedy:
| Device | Output |
|---|---|
| GPU.0 | ")-+&-&)&/-,'/!&*%!'.33).)'03*#*," |
| CPU (251 s) | same string |
| Dummy zeros image | all ! (token id 0) |
Not a GPU plugin bug. CPU matches GPU. Fresh JIT cache, DYNAMIC_QUANTIZATION_GROUP_SIZE=0, official max_new_tokens= kwargs: still garbage.
Community copies (circulus/*-int4-ov, Morteza89/qwen3.8-27b-int4-ov) also ship ignored_scope: null. Do not download those as a fix.
Hub INT8 (OpenVINO/Qwen3.8-27B-int8-ov) is coherent on the same card.
What this repo is
Re-export from BF16:
optimum/ OpenVINO trace ofQwen/Qwen3.8-27B(Transformers 5.2.0 — 5.4.0 refuses the export)- NNCF
compress_weightson the language IR only:INT4_ASYM, group 128, ratio 1.0IgnoredScope(patterns=[".*linear_attn.*"])
convert_tokenizer --with-detokenizer
NNCF bitwidth on the language graph:
| Mode | Share |
|---|---|
| INT4_ASYM g128 | 73% (256 / 545 layers) |
| float (ignored GDN) | 22% (288 / 545) |
| INT8_ASYM per-channel | 5% (1 layer) |
Language bin 21 GB vs official all-INT4 14 GB. The extra 7 GB is GDN left at float.
Runtime
- OpenVINO GenAI 2026.4+ (this IR is a VLM graph). Cascadia 0.2.3 stock
ov-genaiships 2026.2 and SIGSEGVs inTokenizer::setup_tokenizer. - Device:
GPU.0,hint.performance_mode=LATENCY,hint.inference_precision=f16. - Call
pipe.generate(prompt, image=..., generation_config=cfg)by keyword. A positional config is treated as images.
import numpy as np, openvino as ov, openvino_genai as og
from PIL import Image
P = ov.properties
pipe = og.VLMPipeline(
"SergiioB/Qwen3.8-27B-int4-gdn8-ov",
"GPU.0",
**{
P.hint.performance_mode: P.hint.PerformanceMode.LATENCY,
P.hint.inference_precision: ov.Type.f16,
},
)
cfg = og.GenerationConfig()
cfg.max_new_tokens = 128
image = ov.Tensor(np.array(Image.open("photo.jpg").convert("RGB"))[None])
print(pipe.generate("Describe this image.", image=image, generation_config=cfg))
B70 numbers (2026-09-12)
One Arc Pro B70, 150 W cap, GPU.0, greedy. Quality = English on the HF ai2d fire image.
| Engine | Artifact | n=128 wall tok/s | Quality |
|---|---|---|---|
| OpenVINO GenAI 2026.5 | this INT4-GDN8 | 13.89 | PASS |
| OpenVINO GenAI 2026.5 | Hub INT8 | 13.09 | PASS |
| OpenVINO GenAI 2026.5 | Hub INT4 | ~30 (token 0 / punctuation) | FAIL |
Cascadia 0.2.3 qwen35 |
Hub INT8 1-stage shard | 4.56 | PASS (host-side logits) |
| llama.cpp SYCL | Q4_K_M GGUF | 18.58 tg128 | PASS (different artifact) |
Do not put Hub INT4 tok/s on a leaderboard.
License
Apache-2.0, same as Qwen3.8-27B. Intel did not publish these files.