oddadmix/Nawah-BERT-6M-bilingual

🤗 Hugging Face 来源fill-maskapache-2.06M 参数24 MBsafetensors✓ 1 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo oddadmix/Nawah-BERT-6M-bilingual ./model-folder
需要做种者 →

Nawah-BERT-6M-bilingual — pretrained from scratch on 5B Arabic + 5B English tokens

Same architecture as Nawah-BERT-6M-v2 (hidden 128, 8 layers, 2 heads, 5,993,600 params) but pretrained on a genuinely balanced bilingual corpus with a tokenizer built for both languages, rather than an Arabic-only backbone.

Why this exists

A router-head experiment on Nawah-BERT-6M-v2 (Arabic-only pretraining) found that training an English routing head on it alone produced a model that learned nothing — loss stayed flat, every eval landed at random chance. Mixing in Arabic router-training data rescued English performance, but that only fixes the downstream task; the backbone itself was still never taught English. This model tests the more direct fix: pretrain the backbone itself on both languages, from scratch.

What's different from Nawah-BERT-6M-v2

Nawah-BERT-6M-v2 this model
pretraining tokens 15B (Arabic only, 2 stages) 10B (5B Arabic + 5B English, one pass, shuffled)
tokenizer custom_llama_tokenizer (32K, fit on Arabic only) bilingual32k_tokenizer (32K, fit on interleaved Arabic+English documents)
English tokenization cost ~2.02 tokens/word (byte-fragmented) ~1.15 tokens/word
Arabic tokenization cost ~1.20 tokens/word ~1.37 tokens/word (the real, accepted trade-off)
architecture hidden 128, 8 layers identical

The tokenizer swap is not free for Arabic — giving English a fair share of the 32K vocabulary budget costs Arabic some merge efficiency. That trade-off is deliberate and documented in train_tokenizer_bilingual.py.

Training

  • Two 5B-token bins (kaust-generative-ai/fineweb-edu-ar, HuggingFaceFW/fineweb-edu), each tokenized fresh with bilingual32k_tokenizer, sampled 50/50 per chunk with a fixed permutation (DualMemmapDataset, forked from the dialect-encoder run's dual-bin design) — a genuine single pass over both languages with no repetition on either side.
  • 2,048 chunks held out from training entirely (1,019 Arabic / 1,029 English), never seen during pretraining.
  • Same recipe as the rest of the BERT ladder: 30%→15% MLM mask anneal, LR 6e-4 from scratch, 2,048-token context, sparse MLM head (masked positions only).
  • 10B tokens, 38,130 steps, 3h 36m on a single consumer GPU.

Held-out results (2,048 chunks never trained on, measured after training)

MLM loss perplexity masked tokens
Arabic 3.4300 30.88 311,982
English 3.1132 22.49 315,240

Both languages learned; English scoring slightly better than Arabic here despite the tokenizer trade-off running the other way, plausibly reflecting fineweb-edu's more uniform filtering versus fineweb-edu-ar's — not something this run isolates further.

Usage

from transformers import AutoTokenizer, AutoModelForMaskedLM

M = "oddadmix/Nawah-BERT-6M-bilingual"
tok = AutoTokenizer.from_pretrained(M)
model = AutoModelForMaskedLM.from_pretrained(M)

Ships as a plain BertForMaskedLM (the sparse-head training trick used to speed up pretraining is not part of the saved architecture). Downstream fine-tuning (classification, routing) needs mean pooling over non-pad positions, not [CLS] — this is a packed-corpus BERT with no meaningful CLS position, the same trap documented on every other model in this family.

© KAND CA 2026 — PROJECT NAWAH