RikkaBotan/NexteraBERT-Mezzoforte-220M-en

🤗 Hugging Face 来源fill-maskmit213M 参数851 MBsafetensors✓ 8 个校验和今天更新
磁力链接🌱 0✓ 与 Hugging Face 一致

NexteraBERT-Mezzoforte-220M-en

NexteraBERT is a fast, long-context bidirectional encoder. This repository holds the main model of the NexteraBERT paper: an English model with 212.6M parameters, pretrained with masked language modeling (MLM) on 130B tokens.

The encoder has 18 blocks. Each block has a token mixer and a squared-ReLU MLP, both with pre-normalization and a residual connection. The token mixers are:

  • SnowLily (8 blocks): an input-dependent gating liquid mixer derived from liquid time-constant networks. Its cost is linear in input length.
  • NexteraSWA (5 blocks): sliding-window attention with a 256-token window.
  • NexteraSelfAttention (3 blocks): full attention.
  • NexteraHRA (2 blocks): attention over bands of 4 tokens, a cheaper global path.

All attention blocks use grouped-query attention (16 query heads, 8 key/value heads). The full-attention and HRA blocks use SSSMax (Stable Scalable-Softmax), which scales the attention logits with the number of keys so that attention stays selective on long inputs.

Results

(a, b) GLUE and MTEB v2 against pretraining tokens: NexteraBERT at 1.3B, 13B, and 130B tokens, with baselines evaluated under the same protocol. (c) H200 NVL throughput with 65,536 tokens per batch. (d) Exponentiated masked-token loss; the dotted line marks 8,192 tokens, the longest training length, and the percentages give the increase from 8,192 to 65,536 tokens.

Comparison with ModernBERT-base

Both models are fine-tuned and evaluated under the same protocol. NexteraBERT uses 130B pretraining tokens and ModernBERT-base 2T. The better score is in bold.

Benchmark Metric NexteraBERT ModernBERT-base
GLUE mean of 8 tasks 87.90 87.97
MTEB v2 (English, 41 tasks) mean over task types 54.70 53.63
BEIR (15 datasets) nDCG@10 43.89 43.09
NanoBEIR (13 subsets) nDCG@10 55.92 53.76
CodeSearchNet (CoIR) nDCG@10 57.22 56.65
StackOverflowQA nDCG@10 65.71 69.20
MultiLongDocRetrieval (MLDR) nDCG@10 36.25 31.19
Masked-token loss, 8,192 → 65,536 tokens increase (lower is better) +3.8% +83.0%
Throughput at 65,536 tokens (H200 NVL) relative to ModernBERT-base 5.22× 1.00×

GLUE (validation sets)

Task Metric NexteraBERT ModernBERT-base
CoLA MCC 64.64 62.63
SST-2 accuracy 94.07 94.92
MRPC F1 92.81 92.58
STS-B Spearman 91.93 91.98
QQP accuracy 91.58 91.45
MNLI-m accuracy 88.33 88.74
QNLI accuracy 93.50 93.59
RTE accuracy 86.35 87.87
Mean 87.90 87.97

Scores are means over 5 seeds (MRPC, STS-B, RTE), 4 (CoLA), 3 (SST-2), or 1 (QQP, MNLI, QNLI). MRPC, STS-B, RTE, and QNLI start from the MNLI-tuned encoder.

MTEB v2 (English)

Task type NexteraBERT ModernBERT-base
Classification 74.57 73.84
Clustering 43.07 42.12
Pair classification 80.30 79.61
Reranking 41.81 42.02
Retrieval 29.51 26.37
STS 81.65 80.70
Summarization 32.00 30.76
Mean (types) 54.70 53.63
Mean (tasks) 56.77 55.40

Both models are fine-tuned with supervised SimCSE on 312,663 NLI triplets (contrastive loss, maximum length 64) and evaluated at maximum length 512.

Retrieval

Both models are fine-tuned on MS MARCO (1.25M triplets with one hard negative, contrastive loss, mean pooling). BEIR datasets are evaluated at maximum length 512, the code tasks and MLDR at 8,192.

Task NexteraBERT ModernBERT-base
BEIR, mean of 15 datasets 43.89 43.09
CodeSearchNet (CoIR) 57.22 56.65
StackOverflowQA 65.71 69.20
Code, mean of 2 tasks 61.46 62.92
MultiLongDocRetrieval (8,192 tokens) 36.25 31.19
BEIR, per dataset (nDCG@10)
Dataset NexteraBERT ModernBERT-base
TREC-COVID 75.53 75.05
NFCorpus 29.29 26.00
Natural Questions 46.61 45.84
HotpotQA 48.51 48.74
FiQA-2018 31.94 30.66
ArguAna 49.48 46.42
Touche-2020 24.87 23.20
CQADupstack (12 forums) 31.93 32.32
Quora 86.46 86.65
DBPedia 27.86 28.43
SCIDOCS 14.38 13.86
FEVER 66.25 66.59
Climate-FEVER 23.80 22.80
SciFact 60.71 59.97
MS MARCO 40.73 39.85
Mean (15 datasets) 43.89 43.09
NanoBEIR, per subset (nDCG@10)
Subset NexteraBERT ModernBERT-base
MS MARCO 63.52 58.45
NFCorpus 31.34 24.54
Natural Questions 57.92 60.99
HotpotQA 65.28 62.75
FiQA-2018 50.46 41.23
ArguAna 53.10 49.45
Touche-2020 54.53 51.87
Quora 95.56 92.87
DBPedia 49.66 51.01
SCIDOCS 30.86 31.73
FEVER 79.36 80.65
Climate-FEVER 32.21 32.96
SciFact 63.15 60.32
Mean (13 subsets) 55.92 53.76

Usage

The model code is included in this repository, so pass trust_remote_code=True. Tested with transformers 5.12 and PyTorch 2.12.

Masked-token prediction

from transformers import AutoModelForMaskedLM, AutoTokenizer, pipeline

repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True)

fill = pipeline("fill-mask", model=model, tokenizer=tokenizer)
print(fill("The capital of France is [MASK]."))

Hidden states

import torch
from transformers import AutoModel, AutoTokenizer

repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()

batch = tokenizer(["hello world", "a second sentence"], padding=True, return_tensors="pt")
with torch.no_grad():
    hidden = model(**batch).last_hidden_state          # (batch, length, 1024)

mask = batch["attention_mask"].unsqueeze(-1).to(hidden.dtype)
embedding = (hidden * mask).sum(1) / mask.sum(1)       # mean pooling

This is a pretrained encoder, not a sentence-embedding model. Fine-tune it (for example with contrastive learning) before using it for retrieval or similarity.

Unpadded inference

import torch
from transformers import AutoModel, AutoTokenizer

repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, trust_remote_code=True, unpadding=True).cuda().eval()

texts = ["a short query", "a much longer passage about something else entirely"]
batch = tokenizer(texts, padding=True, return_tensors="pt").to("cuda")
with torch.no_grad(), torch.autocast("cuda", dtype=torch.bfloat16):
    hidden = model(**batch).last_hidden_state          # (batch, length, 1024)

Fine-tuning for classification

from transformers import AutoModelForSequenceClassification

model = AutoModelForSequenceClassification.from_pretrained(
    "RikkaBotan/NexteraBERT-Mezzoforte-220M-en",
    trust_remote_code=True,
    num_labels=3,
)

The classification head is newly initialized and must be trained.

Training

Objective MLM, 80/10/10 replacement
Data FineWeb-Edu 90.25%, DCLM 4.75%, StarCoderData 5%
Tokenizer ModernBERT tokenizer (vocabulary 50,368)
Phase 1 120B tokens, inputs up to 1,024 tokens, masking rate 40% → 20% (linear)
Phase 2 10B tokens, inputs up to 8,192 tokens, masking rate 20%, short-context batches on 40% of steps
Optimizer AdamW, β = (0.9, 0.95), weight decay 0.1, gradient clipping 1.0
Peak learning rate 8e-4 (phase 1), 8e-5 (phase 2)
Precision bf16

Architecture

Blocks / width / MLP width 18 / 1,024 / 2,304
Query heads / key-value heads 16 / 8 (head width 64)
Token mixers 8 SnowLily, 5 window attention, 3 full attention, 2 HRA
Window attention window 256
HRA band size 4
RoPE base 100,000 (full attention, HRA); 10,000 (window attention)
Parameters 212.6M (161.1M non-embedding)

Files

  • config.json: model configuration, including the auto_map for trust_remote_code
  • model.safetensors / pytorch_model.bin: encoder weights
  • mlm_head.safetensors: masked-LM prediction head, loaded by AutoModelForMaskedLM
  • tokenizer.json, tokenizer_config.json: tokenizer
  • modeling_nexterabert_hf.py: transformers model classes
  • modeling_nexterabert.py, configuration_nexterabert.py: model and configuration code
  • loading.py, export.py, __init__.py: loading and export helpers
  • assets/summary.png: results figure

🩵 About me 🩵

Japanese independent researcher having shy and pampered personality. Twin-tail hair is a charm point. Interested in nlp. Usually using python and C.

X(Twitter): https://twitter.com/peony__snow