oddadmix/50M-English-MSA-v1

🤗 Hugging Face 来源text-generationapache-2.052M 参数104 MBsafetensors✓ 2 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo oddadmix/50M-English-MSA-v1 ./model-folder
需要做种者 →

50M-English-MSA-v1 — Bidirectional English ↔ MSA (Modern Standard Arabic)

A 51.8M-parameter small language model that translates both ways between English and Modern Standard Arabic (الفصحى). A single set of weights serves both directions; a direction-specific system prompt selects which way to translate.

Finetuned from oddadmix/50M-2048-Emhotob, a tiny Arabic base model trained from scratch.

Evaluation

Evaluated on a deterministic held-out set of 3,000 pairs (seed=42), decoded greedily (do_sample=False, no repetition penalty), scored with sacreBLEU:

Direction sacreBLEU chrF
English → MSA 46.22 66.52
MSA → English 49.85 64.78

Scores are much higher than the dialectal (Egyptian) siblings of this model family: MSA is highly standardized, so a single reference captures most of the valid output space in both directions. The saved weights are the best checkpoint by validation loss (eval_loss=0.4945, epoch 2 of 3).

Decoding note: use plain greedy. A repetition penalty (1.2) was tested across these 50M translation models and lowered BLEU by 5–12 points, because Arabic legitimately repeats short particles that the penalty suppresses.

Example translations

Real greedy-decoded outputs from the held-out set:

English → MSA

English input Model output (MSA)
Thank you very much, you are so kind. شكرًا جزيلًا لك، أنت لطيف جدًا.
May God reward you with goodness for your great hospitality. ليُكافئك الله بالخير لضيافتك العظيمة.
I'm not really scared of anything. أنا لست خائفًا حقًا من أي شيء.

MSA → English

MSA input Model output (English)
شكرًا جزيلًا لك، أنت لطيف للغاية. Thank you so much, you're so sweet.
عزيزتي، المقصد ليس أن لاعبًا واحدًا هو الذي يؤثر على المنتخب الوطني. My dear, the point is not that one player is the one to affect the national team.
أنا لست خائفًا من أي شيء حقًا. I'm not really afraid of anything.

A larger set of 20 examples per direction (with references) is in eval_bidirectional.json.

Usage

ChatML format. Pick the system prompt for the direction you want:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "oddadmix/50M-English-MSA-v1"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16).to("cuda").eval()

SYS_TO_MSA = "أنت مترجم محترف. ترجم النص الإنجليزي إلى اللغة العربية الفصحى."
SYS_TO_EN  = "You are a professional translator. Translate the Modern Standard Arabic text into English."

def translate(text: str, system: str) -> str:
    prompt = (
        f"<|im_start|>system\n{system}<|im_end|>\n"
        f"<|im_start|>user\n{text.strip()}<|im_end|>\n"
        f"<|im_start|>assistant\n"
    )
    ids = tok(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
    if tok.bos_token_id is not None:  # training prepends BOS
        bos = torch.tensor([[tok.bos_token_id]], device=model.device)
        ids["input_ids"] = torch.cat([bos, ids["input_ids"]], dim=1)
        ids["attention_mask"] = torch.cat([torch.ones_like(bos), ids["attention_mask"]], dim=1)
    out = model.generate(**ids, max_new_tokens=256, do_sample=False,
                         eos_token_id=tok.eos_token_id, pad_token_id=tok.pad_token_id)
    return tok.decode(out[0, ids["input_ids"].size(1):], skip_special_tokens=True).strip()

print(translate("Thank you very much, you are so kind.", SYS_TO_MSA))
# → شكرًا جزيلًا لك، أنت لطيف جدًا.
print(translate("أنا لست خائفًا من أي شيء حقًا.", SYS_TO_EN))
# → I'm not really afraid of anything.

Training

  • Base model: oddadmix/50M-2048-Emhotob (Llama arch, ~51.8M params)
  • Dataset: oddadmix/egyptian-msa-2.9-openai-bytedance-translations (132K rows; this model uses the english and msa columns)
  • Method: HuggingFace Trainer, ChatML, prompt-masked cross-entropy (loss only on the assistant turn). Each row is exploded into two training examples (one per direction, ~254K total). Two ChatML special tokens (<|im_start|>, <|im_end|>) were added and embeddings resized.
  • Hyperparameters: 3 epochs · effective batch 64 · LR 3e-4 (cosine, 5% warmup) · bf16 · max length 1024 · load_best_model_at_end on eval_loss.
  • Split: 129,009 train / 3,000 deterministic held-out (seed=42), scored both directions.

Limitations

  • A 50M model: expect errors on rare / technical vocabulary, proper nouns, and occasional drift or truncation on very long inputs. Everyday and conversational text is handled well.
  • Gender is disambiguated only from context; ambiguous inputs may default one way.
  • Trained on the MSA side of an Egyptian-sourced parallel corpus; highly formal, legal, or classical registers are out of scope. For dialectal Egyptian, see the sibling models oddadmix/50M-Egyptian-English-v1 and oddadmix/50M-MSA-Egyptian-v1.

License

Apache-2.0, inherited from the base model.