iapp/ChindaMT-0.8B

🤗 Hugging Face 来源translationapache-2.01.1B 参数2.2 GBsafetensors✓ 3 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo iapp/ChindaMT-0.8B ./model-folder
需要做种者 →

ChindaMT-0.8B

Paper: Reference-Grounded Data Curation for Instruction-Following Thai-English Machine Translation, AACL-IJCNLP 2026 Main Conference

ChindaMT-0.8B is an open-weight Thai-English machine translation model fine-tuned from Qwen/Qwen3.5-0.8B on Grounded, a 1.97M-record dataset built by Reference-Grounded Data Curation (RGDC). It translates in both directions and follows auxiliary rules given in the prompt, such as terminology, register, length, and output format. It is one of three sizes in the ChindaMT family (4B, 2B, 0.8B).

  • Task: Thai-English machine translation with instruction following
  • Base model: Qwen3.5-0.8B
  • Parameters: 0.8B
  • License: Apache-2.0 (inherits the base-model license)

Prompting

Plain translation. Same template for both directions; swap the language line and the source tag:

Translate English to Thai.

EN: The weather is nice today.
Translate Thai to English.

TH: วันนี้อากาศดีมาก

With rules. Add a Rules: block between the language line and the source line. Rules are free-form text:

Translate English to Thai.
Rules:
- Return only the translated text
- Use a clear, professional tone in Thai
- Keep all numerals in Arabic digits

EN: <source text>

Inference

Requires transformers 5.2 or newer:

pip install -U "transformers>=5.2"
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "iapp/ChindaMT-0.8B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16, device_map="auto")

prompt = "Translate English to Thai.\n\nEN: The weather is nice today."
messages = [{"role": "user", "content": prompt}]
text = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=False, enable_thinking=False,
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
out = model.generate(
    **inputs, max_new_tokens=1024, temperature=0.01, top_p=0.7, top_k=20,
    repetition_penalty=1.05, do_sample=True,
)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Thinking mode is disabled (enable_thinking=False); all reported results use this setting.

Training

  • Method: full-parameter supervised fine-tuning with LLaMA-Factory, DeepSpeed ZeRO-2
  • Data: Grounded, 1.97M records
  • Epochs: 1
  • Learning rate: 2e-5, inverse-square-root schedule, 1% warmup
  • Optimizer: AdamW (0.9, 0.999), weight decay 0.01
  • Effective batch size: 64, on two H100 GPUs
  • Cutoff length: 1024 tokens
  • Seed: 42
  • The same recipe trains all three sizes.

Decoding settings

All reported results, and the inference snippet above, use these settings:

  • temperature 0.01, top-p 0.7, top-k 20, repetition penalty 1.05
  • max new tokens 1024
  • thinking mode off (enable_thinking=False)
  • input length: trained with a 1024-token cutoff. Longer inputs worked in our tests but are used at your own risk; recommended: translate long documents paragraph by paragraph, and raise max_new_tokens if an output is cut short.

Evaluation

How to read the tables

  • Metric: length-controlled pairwise win rate (LC%), AlpacaEval-v2 protocol, 400 items per split. Each cell is the LC% of ChindaMT-0.8B against the model named in that row, so values above 50 mean ChindaMT-0.8B wins and 50 is a tie.
  • Judge: Qwen3.6-35B, with two cross-judges from other model families as a check (below).
  • Plain (shared): both systems get the same prompt scaffold, translation only.
  • Plain (own template): the baseline uses its own recommended prompt, the hardest comparison for ChindaMT-0.8B.
  • Constrained: the prompt adds a rules block of one to four constraints; wins here reflect rule-following as well as translation quality.
  • Dashes: omitted because the comparator returns almost no target-language output under the shared prompt.

Against its own base (Qwen3.5-0.8B): Plain 81.5, Constrained 83.5 on CoreEval.

CoreEval (five deployment domains):

ChindaMT-0.8B vs Plain (shared) Plain (own template) Constrained
MiLMMT-46-1B 83.4 50.9 90.3
HY-MT-1.5-1.8B 64.6 63.4 67.8
GemmaX2-28-2B 87.5 51.2 86.5

BroadEval (ten broader domains, cross-domain generalization):

ChindaMT-0.8B vs Plain (shared) Plain (own template) Constrained
MiLMMT-46-1B 92.6 50.2 -
HY-MT-1.5-1.8B 54.4 56.5 61.9
GemmaX2-28-2B - 49.4 -

What the numbers mean

  • Each number is a win rate out of 100: how often the judge preferred the translation from ChindaMT-0.8B over the other model's, on the same 400 sentences, after correcting for output length. 50 is a tie. For example, 64.6 against HY-MT-1.5-1.8B means the judge preferred ChindaMT-0.8B in about 65 of 100 head-to-head comparisons.
  • Under Constrained, a win means the translation was judged better and respected the rules, so those numbers measure rule-following as well as translation quality.
  • In short: ChindaMT-0.8B is preferred over its base and over the 1B to 2B baselines on both splits, with near parity against GemmaX2-28-2B under that model's own template.

External metrics (direction-averaged; quality mean averages CometKiwi, GEMBA-DA, and GEMBA-MQM on a 0-100 scale, higher is better; MetricX-24 is an error score, lower is better):

Benchmark Quality mean MetricX-24 error
FLORES-200 (Wikipedia) 85.9 2.53
WMT24++ en-th (news) 81.2 3.73

Cross-judges. GPT-OSS-120B and Llama-3.3-70B-Instruct agree with the primary judge in direction on every Constrained and Plain (shared) comparison. Full protocol, standard errors, and all baselines are in the paper.

Evaluation suites

Limitations

  • Thai-English only, in both directions.
  • Rule types are those a reference translation can demonstrate: terminology, length, register, and output format. Mandated glossaries and placeholder tokens are outside the training signal by design.
  • The primary judge shares the Qwen family with the model; cross-judges from two other families and non-LLM metrics are reported in the paper as checks.
  • Input length. Trained with a 1024-token cutoff. Inputs above 1k tokens still translated cleanly in our tests, but use them at your own risk. Recommended: translate long documents paragraph by paragraph.

Citation

@misc{chayintr2026rgdc,
  title         = {Reference-Grounded Data Curation for Instruction-Following Thai-English Machine Translation},
  author        = {Chay-intr, Thodsaporn and Harnchang, Krittapad and Thabua, Mahannop and Viriyayudhakorn, Kobkrit and Theeramunkong, Thanaruk},
  year          = {2026},
  eprint        = {2609.34770},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2609.34770},
  note          = {Accepted at AACL-IJCNLP 2026 (Main Conference)}
}

iApp AI Research