oddadmix/Jisr-MT-50M-AllDialects

🤗 Hugging Face sourcetranslationapache-2.048M params193 MBsafetensors✓ 1 checksumupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo oddadmix/Jisr-MT-50M-AllDialects ./model-folder
Needs a seeder →

Jisr-MT-50M-AllDialects

One 49M model, English → 13 Arabic dialects + MSA, selected by tag.

Trained to test whether dialects share enough structure to help each other. They largely do — it beats the single-dialect models on 5 of 9 comparable dialects and ties 2 more, at 1/13th the storage.

Results

UBC-NLP/alexandria test splits, beam 4. "single" is the dedicated 49M model for that dialect.

dialect tag single chrF++ this model Δ chrF++ Δ BLEU
Libyan >>ayl<< 24.17 31.79 +7.62 +6.29
Palestinian >>ajp-ps<< 28.91 33.89 +4.98 +4.87
Syrian >>ajp<< 34.68 37.92 +3.24 +3.06
Saudi >>ars<< 34.50 35.44 +0.94 +1.18
Yemeni >>acq<< 30.12 30.83 +0.71 +0.98
Sudanese >>apd<< 34.11 34.60 +0.49 −0.37
Tunisian >>aeb<< 30.58 30.83 +0.25 −0.42
Lebanese >>apc<< 30.96 28.40 −2.56 −2.89
Moroccan >>ary<< 32.23 30.50 −1.73 −3.66
Egyptian >>arz<< 38.58¹ 37.60 −0.98 −1.32

¹ vs Jisr-MT-50M-Masri-v2, the combined-corpus model.

Sudanese and Tunisian flip sign between the two metrics, so they are genuine ties. The honest tally is 5 clear wins, 3 clear losses, 2 ties.

The pattern is data-size dependent. Gains concentrate in small-corpus dialects — Libyan (12,850 pairs, +7.62) and Palestinian (29,283, +4.98) — and losses in the large ones: Moroccan (155,939) and Egyptian (which also has a second corpus). Dialects with enough data of their own are diluted by sharing capacity; dialects without enough are rescued by it.

MSA retention — the best in the family

model BLEU chrF++
un-finetuned base 30.77 51.09
this model 28.50 49.46
Jisr-MT-50M-Masri-v2 28.21 49.27

It saw 13 dialects' worth of MSA replay rather than one, so >>ara<< received proportionally more gradient.

Compared with a frontier model

gpt-5.1 on the same benchmark (200 turns per dialect, prompted to write dialect not MSA):

dialect this model gpt-5.1 Δ % of gpt-5.1
Moroccan 30.50 20.86 +9.64 146%
Libyan 31.79 34.67 -2.88 92%
Sudanese 34.60 40.14 -5.54 86%
Palestinian 33.89 39.87 -5.98 85%
Tunisian 30.83 38.64 -7.81 80%
Egyptian 37.60 45.59 -7.99 82%
Syrian 37.92 46.16 -8.24 82%
Yemeni 30.83 39.32 -8.49 78%
Lebanese 28.40 37.19 -8.79 76%
Saudi 35.44 48.74 -13.30 73%

gpt-5.1 leads on 9 of 10, by 2.9–13.3 chrF++, averaging 39.21 against this model's ~34. The exception is Moroccan, where gpt-5.1 collapses to 20.86 — less than half its Saudi score — and this model leads by 9.64. See Jisr-MT-50M-Moroccan for that result.

gpt-5.1 also gains +2.3 to +4.8 chrF++ from orthographic normalisation while Jisr models gain ~+1.6, so the raw gaps understate it: trained in-domain, these models already match the references' spelling conventions.

Usage

from transformers import AutoTokenizer, MarianMTModel
tok = AutoTokenizer.from_pretrained("oddadmix/Jisr-MT-50M-AllDialects")
model = MarianMTModel.from_pretrained("oddadmix/Jisr-MT-50M-AllDialects")

def tr(text, tag):
    return tok.decode(model.generate(**tok(f"{tag} {text}", return_tensors="pt"),
                                     num_beams=4)[0], skip_special_tokens=True)

tr("Where are you going?", ">>arz<<")   # Egyptian
tr("Where are you going?", ">>ary<<")   # Moroccan
tr("Where are you going?", ">>ara<<")   # MSA
tag dialect tag dialect
>>arz<< Egyptian >>ars<< Saudi
>>ary<< Moroccan >>afb<< Bahraini (Gulf)
>>arq<< Algerian >>apd<< Sudanese
>>aeb<< Tunisian >>acq<< Yemeni
>>ayl<< Libyan >>acm<< Iraqi
>>apc<< Lebanese >>ajp-ps<< Palestinian
>>ajp<< Syrian >>ara<< MSA

The tag is required — the model cannot infer the target dialect. English → Arabic only.

Note >>apc<< (Lebanese) and >>ajp<< (Syrian) deviate from strict ISO 639-3, where both Levantine variants share apc. They were given distinct tags so the model can separate them.

Training

Base Jisr-MT-50M-tatoeba (49M)
Data 1,458,160 rows across 13 dialects (Egyptian uses the combined corpus)
MSA replay 1×
Budget 1 epoch, bs 64, lr 3e-4 cosine, bf16, 57 min
Vocab 32,014 — 32k base plus 13 dialect tags, each seeded from >>ara<<

Limitations

  • Higher degeneracy than the single models — up to 2.6% of outputs on Sudanese repeat themselves, vs ~1% for the dedicated models. Check outputs before shipping.
  • Algerian, Bahraini and Iraqi are unvalidated. Alexandria has no config for them, so 3 of the 13 tags have no external benchmark at all.
  • Dialect data is Gemma-translated ASR transcripts; for Egyptian, where a purpose-built parallel corpus exists, that corpus alone trains a better model.
  • No standard orthography for any of these dialects — all chrF/BLEU here are depressed by that.
  • English → Arabic only. 49M parameters.