NeoMME-800M predecay checkpoint
Model summary
This repository contains the NeoMME-800M checkpoint at step 450,000 of a 500,000-step pretraining run. As the model name suggests, this checkpoint was saved just before the learning rate decay phase. It is intended as a starting point for continued pretraining with NeoMMEForMaskedLM.
| Item | Value |
|---|---|
| Training step | 450,000 of 500,000 |
| Intended use | Continued pretraining |
| Model class | NeoMMEForMaskedLM |
| Training objective | Predict masked text tokens from visible text and image patches |
[!IMPORTANT]
training_state.ptcontains the original native NeoMME optimizer and training state. It is provided for research and custom continuation workflows, but it is not directly compatible with Hugging Face Trainer or standard AdamW. The file uses NeoMME's custom NorMuon and MasterAdamW state formats, native parameter grouping, and FP32 master weights.
Continue pretraining
The original run used the following optimizer settings:
| Parameter group | Optimizer | Peak learning rate | Weight decay |
|---|---|---|---|
| Two-dimensional weight matrices | NorMuon | 0.010 | 0.01 |
| Embedding tables | AdamW | 0.00075 | 0.01 |
| One-dimensional parameters | AdamW | 0.00075 | 0 |
The original learning-rate schedule used:
- A 300-step linear warmup.
- A stable phase through step 450,000.
- A 50,000-step linear decay from the peak learning rate to 1% of the peak learning rate.
For a Transformers continuation with new optimizers, initialize them at the peak learning rates and apply the 50,000-step linear decay.
License
Model weights are released under the Apache 2.0 license.
Citation
@misc{lac2026neommesingletowermultimodalnativemultilingual,
title={NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference},
author={Aurélien Lac and Tony Wu},
year={2026},
eprint={2609.01657},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2609.01657},
}