OpenThai2.0 - Opensource Thai Knowledge, Document, and Agentic AI (GGUF)
GGUF quantizations of openthai2.0-qwen3.8-27b (v9) for llama.cpp. Includes the MTP head (exported as nextn layers) and the vision projector.
| file | size | use |
|---|---|---|
| openthai2.0-qwen3.8-27b-Q4_K_M.gguf | ~17 GB | recommended balance |
| openthai2.0-qwen3.8-27b-IQ2_M.gguf | ~9.8 GB | recommended 2-bit — imatrix-calibrated (Thai+EN); passes factual sanity checks that plain Q2_K fails |
| openthai2.0-qwen3.8-27b-Q2_K.gguf | ~11 GB | plain 2-bit — ⚠️ factual slips observed; prefer IQ2_M |
| openthai2.0-qwen3.8-27b-Q8_0.gguf | ~29 GB | near-lossless |
| mmproj-openthai2.0-qwen3.8-27b-F16.gguf | — | vision projector (documents/images) |
Run
llama-server -m openthai2.0-qwen3.8-27b-Q4_K_M.gguf --mmproj mmproj-openthai2.0-qwen3.8-27b-F16.gguf -c 32768
⚠️ The model reasons before it answers. Use a large context (32k) and leave max_tokens unset or >= 8192, or replies may come back empty.
Sanity-verified: Q4_K_M (CPU) and IQ2_M (GPU) answer Thai factual prompts correctly; imatrix-a80.dat is included for community re-quants. Full benchmarks and model card: see the main bf16 repo.