NexteraBERT-Mezzoforte-220M-1.3B-en
NexteraBERT is a fast, long-context bidirectional encoder. This repository holds the 1.3B-token model of the NexteraBERT paper: the same 212.6M-parameter English model as NexteraBERT-Mezzoforte-220M-en, pretrained with masked language modeling (MLM) on 1.3B tokens instead of 130B.
The paper trains one architecture at three budgets, 1.3B, 13B, and 130B tokens, to show how quality scales with pretraining. Each budget is a separate run of the same recipe, not an intermediate checkpoint of the 130B run. At 1.3B tokens, NexteraBERT scores 0.42 points above OptiBERTneo trained on the same number of tokens on MTEB v2, at 30% lower estimated compute. For downstream use, the 130B model NexteraBERT-Mezzoforte-220M-en scores higher on every GLUE task and every MTEB task type.
The encoder has 18 blocks. Each block has a token mixer and a squared-ReLU MLP, both with pre-normalization and a residual connection. The token mixers are:
- SnowLily (8 blocks): an input-dependent gating liquid mixer derived from liquid time-constant networks. Its cost is linear in input length.
- NexteraSWA (5 blocks): sliding-window attention with a 256-token window.
- NexteraSelfAttention (3 blocks): full attention.
- NexteraHRA (2 blocks): attention over bands of 4 tokens, a cheaper global path.
All attention blocks use grouped-query attention (16 query heads, 8 key/value heads). The full-attention and HRA blocks use SSSMax (Stable Scalable-Softmax), which scales the attention logits with the number of keys so that attention stays selective on long inputs.
Results
(a, b) GLUE and MTEB v2 against pretraining tokens: NexteraBERT at 1.3B (this model), 13B, and 130B tokens, with ModernBERT-base (2T) and NeoBERT (2.1T) evaluated under the same protocol; OptiBERTneo values are published. (c) H200 NVL throughput with 65,536 tokens per batch. (d) Exponentiated masked-token loss; the dotted line marks 8,192 tokens, the longest training length, and the percentages give the increase from 8,192 to 65,536 tokens. Panels (c) and (d) show the 130B model. Throughput depends only on the architecture, which this model shares; length extrapolation was not measured at 1.3B tokens.
Comparison across pretraining budgets
All three budgets are fine-tuned and evaluated under the same protocol.
| Benchmark | Metric | 1.3B (this model) | 13B | 130B |
|---|---|---|---|---|
| GLUE | mean of 8 tasks | 77.54 | 85.45 | 87.90 |
| MTEB v2 (English, 41 tasks) | mean over task types | 47.82 | 51.92 | 54.70 |
Retrieval (BEIR, NanoBEIR, code, and MLDR) and length extrapolation were evaluated for the 130B model only; see NexteraBERT-Mezzoforte-220M-en.
Comparison with OptiBERTneo at 1.3B tokens
OptiBERTneo (Dervishi et al., 2025) is a standard Transformer encoder with the NeoBERT architecture (198M non-embedding parameters), trained on FineWeb-Edu with 1,024-token inputs at the same token budgets as NexteraBERT; 90.25% of NexteraBERT's data comes from the same corpus. Its scores are published values, not re-evaluated under our protocol: they come from another tokenizer and the fine-tuning runs of that study, so the comparison rules out the data as the explanation but is not fully controlled. The better score is in bold.
| Benchmark | Metric | NexteraBERT | OptiBERTneo |
|---|---|---|---|
| GLUE | mean of 8 tasks | 77.54 | 76.0 |
| MTEB v2 (English, 41 tasks) | mean over task types | 47.82 | 47.40 |
| Estimated pretraining compute | FLOPs (lower is better) | 1.33 × 10¹⁸ | 1.9 × 10¹⁸ |
Compute is estimated with the OptiBERT convention (six FLOPs per non-embedding parameter per token, plus the pairwise attention terms at 1,024 tokens), not measured.
GLUE (validation sets)
| Task | Metric | 1.3B (this model) | 13B | 130B |
|---|---|---|---|---|
| CoLA | MCC | 39.59 | 61.33 | 64.64 |
| SST-2 | accuracy | 86.77 | 92.58 | 94.07 |
| MRPC | F1 | 88.42 | 91.15 | 92.81 |
| STS-B | Spearman | 85.16 | 90.30 | 91.93 |
| QQP | accuracy | 89.22 | 91.12 | 91.58 |
| MNLI-m | accuracy | 76.79 | 84.95 | 88.33 |
| QNLI | accuracy | 84.99 | 91.36 | 93.50 |
| RTE | accuracy | 69.39 | 80.79 | 86.35 |
| Mean | 77.54 | 85.45 | 87.90 |
Scores are means over 5 seeds (MRPC, STS-B, RTE), 4 (CoLA), 3 (SST-2), or 1 (QQP, MNLI, QNLI). MRPC, STS-B, RTE, and QNLI start from the MNLI-tuned encoder.
MTEB v2 (English)
| Task type | 1.3B (this model) | 13B | 130B |
|---|---|---|---|
| Classification | 64.66 | 71.39 | 74.57 |
| Clustering | 39.64 | 42.31 | 43.07 |
| Pair classification | 67.36 | 74.92 | 80.30 |
| Reranking | 40.06 | 41.26 | 41.81 |
| Retrieval | 22.07 | 23.22 | 29.51 |
| STS | 73.13 | 78.96 | 81.65 |
| Summarization | 27.81 | 31.37 | 32.00 |
| Mean (types) | 47.82 | 51.92 | 54.70 |
| Mean (tasks) | 49.35 | 53.44 | 56.77 |
Each model is fine-tuned with supervised SimCSE on 312,663 NLI triplets (contrastive loss, maximum length 64) and evaluated at maximum length 512.
Usage
The model code is included in this repository, so pass trust_remote_code=True.
Tested with transformers 5.12 and PyTorch 2.12.
Masked-token prediction
from transformers import AutoModelForMaskedLM, AutoTokenizer, pipeline
repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-1.3B-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True)
fill = pipeline("fill-mask", model=model, tokenizer=tokenizer)
print(fill("The capital of France is [MASK]."))
Hidden states
import torch
from transformers import AutoModel, AutoTokenizer
repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-1.3B-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()
batch = tokenizer(["hello world", "a second sentence"], padding=True, return_tensors="pt")
with torch.no_grad():
hidden = model(**batch).last_hidden_state # (batch, length, 1024)
mask = batch["attention_mask"].unsqueeze(-1).to(hidden.dtype)
embedding = (hidden * mask).sum(1) / mask.sum(1) # mean pooling
This is a pretrained encoder, not a sentence-embedding model. Fine-tune it (for example with contrastive learning) before using it for retrieval or similarity.
Unpadded inference
import torch
from transformers import AutoModel, AutoTokenizer
repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-1.3B-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, trust_remote_code=True, unpadding=True).cuda().eval()
texts = ["a short query", "a much longer passage about something else entirely"]
batch = tokenizer(texts, padding=True, return_tensors="pt").to("cuda")
with torch.no_grad(), torch.autocast("cuda", dtype=torch.bfloat16):
hidden = model(**batch).last_hidden_state # (batch, length, 1024)
Fine-tuning for classification
from transformers import AutoModelForSequenceClassification
model = AutoModelForSequenceClassification.from_pretrained(
"RikkaBotan/NexteraBERT-Mezzoforte-220M-1.3B-en",
trust_remote_code=True,
num_labels=3,
)
The classification head is newly initialized and must be trained.
Training
| Objective | MLM, 80/10/10 replacement |
| Data | FineWeb-Edu 90.25%, DCLM 4.75%, StarCoderData 5% |
| Tokenizer | ModernBERT tokenizer (vocabulary 50,368) |
| Phase 1 | 1.2B tokens, inputs up to 1,024 tokens, masking rate 40% → 20% (linear) |
| Phase 2 | 0.1B tokens, inputs up to 8,192 tokens, masking rate 20%, short-context batches on 40% of steps |
| Optimizer | AdamW, β = (0.9, 0.95), weight decay 0.1, gradient clipping 1.0 |
| Peak learning rate | 8e-4 (phase 1), 8e-5 (phase 2) |
| Precision | bf16 |
| MLM training loss | 2.16 (end of phase 1), 2.10 (end of phase 2) |
The recipe is the same at every budget. The masking-rate and learning-rate schedules (linear warm-up, cosine decay) span each phase of this run, so they are shorter than in the 130B run.
Architecture
| Blocks / width / MLP width | 18 / 1,024 / 2,304 |
| Query heads / key-value heads | 16 / 8 (head width 64) |
| Token mixers | 8 SnowLily, 5 window attention, 3 full attention, 2 HRA |
| Window attention | window 256 |
| HRA | band size 4 |
| RoPE base | 100,000 (full attention, HRA); 10,000 (window attention) |
| Parameters | 212.6M (161.1M non-embedding) |
Files
config.json: model configuration, including theauto_mapfortrust_remote_codemodel.safetensors/pytorch_model.bin: encoder weightsmlm_head.safetensors: masked-LM prediction head, loaded byAutoModelForMaskedLMtokenizer.json,tokenizer_config.json: tokenizermodeling_nexterabert_hf.py:transformersmodel classesmodeling_nexterabert.py,configuration_nexterabert.py: model and configuration codeloading.py,export.py,__init__.py: loading and export helpers
🩵 About me 🩵
Japanese independent researcher having shy and pampered personality. Twin-tail hair is a charm point. Interested in nlp. Usually using python and C.
X(Twitter): https://twitter.com/peony__snow