NexteraBERT-Mezzoforte-220M-en
NexteraBERT is a fast, long-context bidirectional encoder. This repository holds the main model of the NexteraBERT paper: an English model with 212.6M parameters, pretrained with masked language modeling (MLM) on 130B tokens.
The encoder has 18 blocks. Each block has a token mixer and a squared-ReLU MLP, both with pre-normalization and a residual connection. The token mixers are:
- SnowLily (8 blocks): an input-dependent gating liquid mixer derived from liquid time-constant networks. Its cost is linear in input length.
- NexteraSWA (5 blocks): sliding-window attention with a 256-token window.
- NexteraSelfAttention (3 blocks): full attention.
- NexteraHRA (2 blocks): attention over bands of 4 tokens, a cheaper global path.
All attention blocks use grouped-query attention (16 query heads, 8 key/value heads). The full-attention and HRA blocks use SSSMax (Stable Scalable-Softmax), which scales the attention logits with the number of keys so that attention stays selective on long inputs.
Results
(a, b) GLUE and MTEB v2 against pretraining tokens: NexteraBERT at 1.3B, 13B, and 130B tokens, with baselines evaluated under the same protocol. (c) H200 NVL throughput with 65,536 tokens per batch. (d) Exponentiated masked-token loss; the dotted line marks 8,192 tokens, the longest training length, and the percentages give the increase from 8,192 to 65,536 tokens.
Comparison with ModernBERT-base
Both models are fine-tuned and evaluated under the same protocol. NexteraBERT uses 130B pretraining tokens and ModernBERT-base 2T. The better score is in bold.
| Benchmark | Metric | NexteraBERT | ModernBERT-base |
|---|---|---|---|
| GLUE | mean of 8 tasks | 87.90 | 87.97 |
| MTEB v2 (English, 41 tasks) | mean over task types | 54.70 | 53.63 |
| BEIR (15 datasets) | nDCG@10 | 43.89 | 43.09 |
| NanoBEIR (13 subsets) | nDCG@10 | 55.92 | 53.76 |
| CodeSearchNet (CoIR) | nDCG@10 | 57.22 | 56.65 |
| StackOverflowQA | nDCG@10 | 65.71 | 69.20 |
| MultiLongDocRetrieval (MLDR) | nDCG@10 | 36.25 | 31.19 |
| Masked-token loss, 8,192 → 65,536 tokens | increase (lower is better) | +3.8% | +83.0% |
| Throughput at 65,536 tokens (H200 NVL) | relative to ModernBERT-base | 5.22× | 1.00× |
GLUE (validation sets)
| Task | Metric | NexteraBERT | ModernBERT-base |
|---|---|---|---|
| CoLA | MCC | 64.64 | 62.63 |
| SST-2 | accuracy | 94.07 | 94.92 |
| MRPC | F1 | 92.81 | 92.58 |
| STS-B | Spearman | 91.93 | 91.98 |
| QQP | accuracy | 91.58 | 91.45 |
| MNLI-m | accuracy | 88.33 | 88.74 |
| QNLI | accuracy | 93.50 | 93.59 |
| RTE | accuracy | 86.35 | 87.87 |
| Mean | 87.90 | 87.97 |
Scores are means over 5 seeds (MRPC, STS-B, RTE), 4 (CoLA), 3 (SST-2), or 1 (QQP, MNLI, QNLI). MRPC, STS-B, RTE, and QNLI start from the MNLI-tuned encoder.
MTEB v2 (English)
| Task type | NexteraBERT | ModernBERT-base |
|---|---|---|
| Classification | 74.57 | 73.84 |
| Clustering | 43.07 | 42.12 |
| Pair classification | 80.30 | 79.61 |
| Reranking | 41.81 | 42.02 |
| Retrieval | 29.51 | 26.37 |
| STS | 81.65 | 80.70 |
| Summarization | 32.00 | 30.76 |
| Mean (types) | 54.70 | 53.63 |
| Mean (tasks) | 56.77 | 55.40 |
Both models are fine-tuned with supervised SimCSE on 312,663 NLI triplets (contrastive loss, maximum length 64) and evaluated at maximum length 512.
Retrieval
Both models are fine-tuned on MS MARCO (1.25M triplets with one hard negative, contrastive loss, mean pooling). BEIR datasets are evaluated at maximum length 512, the code tasks and MLDR at 8,192.
| Task | NexteraBERT | ModernBERT-base |
|---|---|---|
| BEIR, mean of 15 datasets | 43.89 | 43.09 |
| CodeSearchNet (CoIR) | 57.22 | 56.65 |
| StackOverflowQA | 65.71 | 69.20 |
| Code, mean of 2 tasks | 61.46 | 62.92 |
| MultiLongDocRetrieval (8,192 tokens) | 36.25 | 31.19 |
| Dataset | NexteraBERT | ModernBERT-base |
|---|---|---|
| TREC-COVID | 75.53 | 75.05 |
| NFCorpus | 29.29 | 26.00 |
| Natural Questions | 46.61 | 45.84 |
| HotpotQA | 48.51 | 48.74 |
| FiQA-2018 | 31.94 | 30.66 |
| ArguAna | 49.48 | 46.42 |
| Touche-2020 | 24.87 | 23.20 |
| CQADupstack (12 forums) | 31.93 | 32.32 |
| Quora | 86.46 | 86.65 |
| DBPedia | 27.86 | 28.43 |
| SCIDOCS | 14.38 | 13.86 |
| FEVER | 66.25 | 66.59 |
| Climate-FEVER | 23.80 | 22.80 |
| SciFact | 60.71 | 59.97 |
| MS MARCO | 40.73 | 39.85 |
| Mean (15 datasets) | 43.89 | 43.09 |
| Subset | NexteraBERT | ModernBERT-base |
|---|---|---|
| MS MARCO | 63.52 | 58.45 |
| NFCorpus | 31.34 | 24.54 |
| Natural Questions | 57.92 | 60.99 |
| HotpotQA | 65.28 | 62.75 |
| FiQA-2018 | 50.46 | 41.23 |
| ArguAna | 53.10 | 49.45 |
| Touche-2020 | 54.53 | 51.87 |
| Quora | 95.56 | 92.87 |
| DBPedia | 49.66 | 51.01 |
| SCIDOCS | 30.86 | 31.73 |
| FEVER | 79.36 | 80.65 |
| Climate-FEVER | 32.21 | 32.96 |
| SciFact | 63.15 | 60.32 |
| Mean (13 subsets) | 55.92 | 53.76 |
Usage
The model code is included in this repository, so pass trust_remote_code=True.
Tested with transformers 5.12 and PyTorch 2.12.
Masked-token prediction
from transformers import AutoModelForMaskedLM, AutoTokenizer, pipeline
repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True)
fill = pipeline("fill-mask", model=model, tokenizer=tokenizer)
print(fill("The capital of France is [MASK]."))
Hidden states
import torch
from transformers import AutoModel, AutoTokenizer
repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()
batch = tokenizer(["hello world", "a second sentence"], padding=True, return_tensors="pt")
with torch.no_grad():
hidden = model(**batch).last_hidden_state # (batch, length, 1024)
mask = batch["attention_mask"].unsqueeze(-1).to(hidden.dtype)
embedding = (hidden * mask).sum(1) / mask.sum(1) # mean pooling
This is a pretrained encoder, not a sentence-embedding model. Fine-tune it (for example with contrastive learning) before using it for retrieval or similarity.
Unpadded inference
import torch
from transformers import AutoModel, AutoTokenizer
repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, trust_remote_code=True, unpadding=True).cuda().eval()
texts = ["a short query", "a much longer passage about something else entirely"]
batch = tokenizer(texts, padding=True, return_tensors="pt").to("cuda")
with torch.no_grad(), torch.autocast("cuda", dtype=torch.bfloat16):
hidden = model(**batch).last_hidden_state # (batch, length, 1024)
Fine-tuning for classification
from transformers import AutoModelForSequenceClassification
model = AutoModelForSequenceClassification.from_pretrained(
"RikkaBotan/NexteraBERT-Mezzoforte-220M-en",
trust_remote_code=True,
num_labels=3,
)
The classification head is newly initialized and must be trained.
Training
| Objective | MLM, 80/10/10 replacement |
| Data | FineWeb-Edu 90.25%, DCLM 4.75%, StarCoderData 5% |
| Tokenizer | ModernBERT tokenizer (vocabulary 50,368) |
| Phase 1 | 120B tokens, inputs up to 1,024 tokens, masking rate 40% → 20% (linear) |
| Phase 2 | 10B tokens, inputs up to 8,192 tokens, masking rate 20%, short-context batches on 40% of steps |
| Optimizer | AdamW, β = (0.9, 0.95), weight decay 0.1, gradient clipping 1.0 |
| Peak learning rate | 8e-4 (phase 1), 8e-5 (phase 2) |
| Precision | bf16 |
Architecture
| Blocks / width / MLP width | 18 / 1,024 / 2,304 |
| Query heads / key-value heads | 16 / 8 (head width 64) |
| Token mixers | 8 SnowLily, 5 window attention, 3 full attention, 2 HRA |
| Window attention | window 256 |
| HRA | band size 4 |
| RoPE base | 100,000 (full attention, HRA); 10,000 (window attention) |
| Parameters | 212.6M (161.1M non-embedding) |
Files
config.json: model configuration, including theauto_mapfortrust_remote_codemodel.safetensors/pytorch_model.bin: encoder weightsmlm_head.safetensors: masked-LM prediction head, loaded byAutoModelForMaskedLMtokenizer.json,tokenizer_config.json: tokenizermodeling_nexterabert_hf.py:transformersmodel classesmodeling_nexterabert.py,configuration_nexterabert.py: model and configuration codeloading.py,export.py,__init__.py: loading and export helpersassets/summary.png: results figure
🩵 About me 🩵
Japanese independent researcher having shy and pampered personality. Twin-tail hair is a charm point. Interested in nlp. Usually using python and C.
X(Twitter): https://twitter.com/peony__snow