Qwen3.8-Distilled-2B-NPU2-NPU2
OpenFlowLM Q4NX conversion of empero-ai/Qwen3.8-2B-Distill for AMD XDNA NPU inference.
This repository contains a quantized Q4NX port of the model, compiled for the OpenFlowLM (OFLM) runtime. It is not a GGUF file.
| Item | Value |
|---|---|
| Source model | Aempero-ai/Qwen3.8-2B-Distill |
| Source GGUF | Qwen3.8-2B-Distill-Q8_0.gguf |
| Weights | model.q4nx (2.29 GB) |
| Modality | language / vision |
| OFLM version | 0.1.0 |
| Converted | 2026-09-27 |
Install and run
This repository works with oflm-add, a small installer that copies the model
into the OpenFlowLM user directory and registers the tag. It never
modifies the system OpenFlowLM install.
pip install oflm-add or uv tool install oflm-add
uv tool install oflm-add
oflm-add Atomic-Germ/Qwen3.8-Distilled-2B-NPU2-NPU2 --family qwen3.5 --xclbin-from Qwen3.8-Distilled-2B-NPU2-NPU2
OFLM_CONFIG_PATH="$HOME/.config/oflm/model_list.json" OFLM_XCLBIN_PATH="$HOME/.config/oflm" oflm run Qwen3.8-Distilled-2B-NPU2-NPU2
Files
| File | Description |
|---|---|
model.q4nx |
Quantized weights (Q8_0 / Q4_1 / BF16) |
config.json |
OFLM runtime configuration |
tokenizer.json |
Tokenizer vocabulary |
tokenizer_config.json |
Tokenizer configuration |
chat_template.jinja |
Chat template |
vision_weight.q4nx |
Vision model |
Source model card
See the original model card: Atomic-Germ/Qwen3.8-2B-Distill-NPU2