RikkaBotan/NexteraBERT-Mezzoforte-220M-13B-en

🤗 Hugging Face sourcefill-maskmit213M params851 MBsafetensors✓ 3 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo RikkaBotan/NexteraBERT-Mezzoforte-220M-13B-en ./model-folder
Needs a seeder →

NexteraBERT-Mezzoforte-220M-13B-en

NexteraBERT is a fast, long-context bidirectional encoder. This repository holds the 13B-token model of the NexteraBERT paper: the same 212.6M-parameter English model as NexteraBERT-Mezzoforte-220M-en, pretrained with masked language modeling (MLM) on 13B tokens instead of 130B.

The paper trains one architecture at three budgets, 1.3B, 13B, and 130B tokens, to show how quality scales with pretraining. Each budget is a separate run of the same recipe, not an intermediate checkpoint of the 130B run. At 13B tokens, NexteraBERT exceeds the best compute-optimal OptiBERT model on MTEB v2 at 5.27 times lower estimated compute. For downstream use, the 130B model NexteraBERT-Mezzoforte-220M-en scores higher on every GLUE task and every MTEB task type.

The encoder has 18 blocks. Each block has a token mixer and a squared-ReLU MLP, both with pre-normalization and a residual connection. The token mixers are:

  • SnowLily (8 blocks): an input-dependent gating liquid mixer derived from liquid time-constant networks. Its cost is linear in input length.
  • NexteraSWA (5 blocks): sliding-window attention with a 256-token window.
  • NexteraSelfAttention (3 blocks): full attention.
  • NexteraHRA (2 blocks): attention over bands of 4 tokens, a cheaper global path.

All attention blocks use grouped-query attention (16 query heads, 8 key/value heads). The full-attention and HRA blocks use SSSMax (Stable Scalable-Softmax), which scales the attention logits with the number of keys so that attention stays selective on long inputs.

Results

(a, b) GLUE and MTEB v2 against pretraining tokens: NexteraBERT at 1.3B, 13B (this model), and 130B tokens, with ModernBERT-base (2T) and NeoBERT (2.1T) evaluated under the same protocol; OptiBERTneo values are published. (c) H200 NVL throughput with 65,536 tokens per batch. (d) Exponentiated masked-token loss; the dotted line marks 8,192 tokens, the longest training length, and the percentages give the increase from 8,192 to 65,536 tokens. Panels (c) and (d) show the 130B model. Throughput depends only on the architecture, which this model shares; length extrapolation was not measured at 13B tokens.

Comparison across pretraining budgets

All three budgets are fine-tuned and evaluated under the same protocol.

Benchmark Metric 1.3B 13B (this model) 130B
GLUE mean of 8 tasks 77.54 85.45 87.90
MTEB v2 (English, 41 tasks) mean over task types 47.82 51.92 54.70

Retrieval (BEIR, NanoBEIR, code, and MLDR) and length extrapolation were evaluated for the 130B model only; see NexteraBERT-Mezzoforte-220M-en.

Comparison with OptiBERTneo at 13B tokens

OptiBERTneo (Dervishi et al., 2025) is a standard Transformer encoder with the NeoBERT architecture (198M non-embedding parameters), trained on FineWeb-Edu with 1,024-token inputs at the same token budgets as NexteraBERT; 90.25% of NexteraBERT's data comes from the same corpus. Its scores are published values, not re-evaluated under our protocol: they come from another tokenizer and the fine-tuning runs of that study, so the comparison rules out the data as the explanation but is not fully controlled. The better score is in bold.

Benchmark Metric NexteraBERT OptiBERTneo
GLUE mean of 8 tasks 85.45 84.6
MTEB v2 (English, 41 tasks) mean over task types 51.92 50.50
Estimated pretraining compute FLOPs (lower is better) 1.33 × 10¹⁹ 1.9 × 10¹⁹

Compute is estimated with the OptiBERT convention (six FLOPs per non-embedding parameter per token, plus the pairwise attention terms at 1,024 tokens), not measured. The best compute-optimal OptiBERT model reaches 51.60 on MTEB v2 at 7.0 × 10¹⁹ FLOPs; this model exceeds it at 5.27 times lower estimated compute.

GLUE (validation sets)

Task Metric 1.3B 13B (this model) 130B
CoLA MCC 39.59 61.33 64.64
SST-2 accuracy 86.77 92.58 94.07
MRPC F1 88.42 91.15 92.81
STS-B Spearman 85.16 90.30 91.93
QQP accuracy 89.22 91.12 91.58
MNLI-m accuracy 76.79 84.95 88.33
QNLI accuracy 84.99 91.36 93.50
RTE accuracy 69.39 80.79 86.35
Mean 77.54 85.45 87.90

Scores are means over 5 seeds (MRPC, STS-B, RTE), 4 (CoLA), 3 (SST-2), or 1 (QQP, MNLI, QNLI). MRPC, STS-B, RTE, and QNLI start from the MNLI-tuned encoder.

MTEB v2 (English)

Task type 1.3B 13B (this model) 130B
Classification 64.66 71.39 74.57
Clustering 39.64 42.31 43.07
Pair classification 67.36 74.92 80.30
Reranking 40.06 41.26 41.81
Retrieval 22.07 23.22 29.51
STS 73.13 78.96 81.65
Summarization 27.81 31.37 32.00
Mean (types) 47.82 51.92 54.70
Mean (tasks) 49.35 53.44 56.77

Each model is fine-tuned with supervised SimCSE on 312,663 NLI triplets (contrastive loss, maximum length 64) and evaluated at maximum length 512.

Usage

The model code is included in this repository, so pass trust_remote_code=True. Tested with transformers 5.12 and PyTorch 2.12.

Masked-token prediction

from transformers import AutoModelForMaskedLM, AutoTokenizer, pipeline

repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-13B-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True)

fill = pipeline("fill-mask", model=model, tokenizer=tokenizer)
print(fill("The capital of France is [MASK]."))

Hidden states

import torch
from transformers import AutoModel, AutoTokenizer

repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-13B-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()

batch = tokenizer(["hello world", "a second sentence"], padding=True, return_tensors="pt")
with torch.no_grad():
    hidden = model(**batch).last_hidden_state          # (batch, length, 1024)

mask = batch["attention_mask"].unsqueeze(-1).to(hidden.dtype)
embedding = (hidden * mask).sum(1) / mask.sum(1)       # mean pooling

This is a pretrained encoder, not a sentence-embedding model. Fine-tune it (for example with contrastive learning) before using it for retrieval or similarity.

Unpadded inference

import torch
from transformers import AutoModel, AutoTokenizer

repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-13B-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, trust_remote_code=True, unpadding=True).cuda().eval()

texts = ["a short query", "a much longer passage about something else entirely"]
batch = tokenizer(texts, padding=True, return_tensors="pt").to("cuda")
with torch.no_grad(), torch.autocast("cuda", dtype=torch.bfloat16):
    hidden = model(**batch).last_hidden_state          # (batch, length, 1024)

Fine-tuning for classification

from transformers import AutoModelForSequenceClassification

model = AutoModelForSequenceClassification.from_pretrained(
    "RikkaBotan/NexteraBERT-Mezzoforte-220M-13B-en",
    trust_remote_code=True,
    num_labels=3,
)

The classification head is newly initialized and must be trained.

Training

Objective MLM, 80/10/10 replacement
Data FineWeb-Edu 90.25%, DCLM 4.75%, StarCoderData 5%
Tokenizer ModernBERT tokenizer (vocabulary 50,368)
Phase 1 12B tokens, inputs up to 1,024 tokens, masking rate 40% → 20% (linear)
Phase 2 1B tokens, inputs up to 8,192 tokens, masking rate 20%, short-context batches on 40% of steps
Optimizer AdamW, β = (0.9, 0.95), weight decay 0.1, gradient clipping 1.0
Peak learning rate 8e-4 (phase 1), 8e-5 (phase 2)
Precision bf16
MLM training loss 1.31 (end of phase 1), 1.28 (end of phase 2)

The recipe is the same at every budget. The masking-rate and learning-rate schedules (linear warm-up, cosine decay) span each phase of this run, so they are shorter than in the 130B run.

Architecture

Blocks / width / MLP width 18 / 1,024 / 2,304
Query heads / key-value heads 16 / 8 (head width 64)
Token mixers 8 SnowLily, 5 window attention, 3 full attention, 2 HRA
Window attention window 256
HRA band size 4
RoPE base 100,000 (full attention, HRA); 10,000 (window attention)
Parameters 212.6M (161.1M non-embedding)

Files

  • config.json: model configuration, including the auto_map for trust_remote_code
  • model.safetensors / pytorch_model.bin: encoder weights
  • mlm_head.safetensors: masked-LM prediction head, loaded by AutoModelForMaskedLM
  • tokenizer.json, tokenizer_config.json: tokenizer
  • modeling_nexterabert_hf.py: transformers model classes
  • modeling_nexterabert.py, configuration_nexterabert.py: model and configuration code
  • loading.py, export.py, __init__.py: loading and export helpers

🩵 About me 🩵

Japanese independent researcher having shy and pampered personality. Twin-tail hair is a charm point. Interested in nlp. Usually using python and C.

X(Twitter): https://twitter.com/peony__snow