iapp/OpenThai-SystemOne-MLX-8bit

🤗 Hugging Face sourcetext-classificationapache-2.0752M params1.5 GBsafetensors✓ 3 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo iapp/OpenThai-SystemOne-MLX-8bit ./model-folder
Needs a seeder →

OpenThai-SystemOne — mlx-8bit

OpenThai-SystemOne is an open Thai + English System One decision model: one forward pass answers typed questions (choice over up to 255 options, ordinal score, yes/no noul) about a text / JSON state with calibrated probabilities, no text generation. It is a Qwen3.5-0.8B text tower (Thai continued pre-training) plus a 256-slot decision head. This repo is a quantization of v0.3 (commit f3709948).

What is quantized: the tower including the token embeddings (MLX quantizes the embedding table too). The 256-slot decision head and the per-type temperatures stay in fp32 (head.safetensors). Quantization therefore only perturbs the hidden state the head reads.

Format: MLX 8-bit affine (group 64) for Apple Silicon (mlx-lm). mlx-lm runs the tower; the included client applies the decision head on the final hidden states. Size: 800 MB.

Measured on a MacBook Pro M3 Max: 4-bit ≈ 19 ms per 3-question Thai decision (the PyTorch model on MPS: ~150 ms).

Usage

pip install mlx-lm torch transformers safetensors pydantic
huggingface-cli download iapp/OpenThai-SystemOne-MLX-8bit --local-dir openthai-mlx
import sys; sys.path.insert(0, "openthai-mlx")
from openthai_systemone.mlx_client import MLXSystemOneClient
c = MLXSystemOneClient("openthai-mlx")
r = c.system_one("ร้านนี้อาหารอร่อยมาก แต่รอนานเกือบชั่วโมง", {"sentiment": {"type": "choice", "instructions": "ความรู้สึก",
                 "criteria": {"บวก": None, "ลบ": None, "กลาง": None}}})
print(r.answers["sentiment"].choice, r.answers["sentiment"].probabilities)

Accuracy vs the bf16 original (same records, single option order, first 800 per set)

Macro: public 74.2 (original 74.3), Thai 80.0 (original 80.1).

subset bf16 original this Δ
public 13-subset bench
aegis2 (noul) 83.2 83.2 +0.0
boolq (noul) 79.7 79.3 -0.3
civil_comments (noul) 79.0 79.0 +0.0
helpsteer2 (score) 41.6 41.6 +0.0
massive-de-DE (choice) 88.3 88.3 +0.0
massive-en-US (choice) 88.3 88.3 +0.0
multinli (choice) 89.0 88.6 -0.3
paws (noul) 94.0 94.0 +0.0
pubmedqa (choice) 64.0 64.0 +0.0
squad2 (noul) 89.3 89.3 +0.0
summeval-consistency (score) 75.0 75.7 +0.7
summeval-relevance (score) 21.7 21.2 -0.4
vitaminc-dev (choice) 72.5 72.0 -0.5
macro, public 13-subset bench 74.3 74.2 -0.1
Thai held-out / eval sets
banking77 (choice) 59.1 59.1 +0.0
contrastive_th (choice) 80.7 80.7 +0.0
contrastive_th (noul) 83.5 83.5 +0.0
contrastive_th (score) 78.6 76.8 -1.8
massive_th (choice) 90.6 91.0 +0.4
prachathai (choice) 98.3 98.3 +0.0
prachathai (noul) 93.4 93.5 +0.1
sib200_th (choice) 77.9 78.4 +0.5
wisesight (choice) 48.9 49.0 +0.1
wongnai (score) 64.5 64.8 +0.2
xlam_tools (choice) 99.4 99.4 +0.0
xnli_th (choice) 79.8 79.8 +0.0
xnli_th (noul) 86.8 86.4 -0.4
macro, Thai held-out / eval sets 80.1 80.0 -0.1

Notes

  • Scores are single-option-order accuracy on the first 800 records of each set (scripts/06_eval.py --limit 800), the same records for the original and the quantization. score subsets report exact level accuracy.
  • Base model, data, training and the full benchmark tables: iapp/OpenThai-SystemOne.
  • License Apache-2.0 (same as the base). Built by iApp Technology / OpenThaiGPT.