💻 GitHub · 🤗 Hugging Face · 🌐 Website
Nex-N2.5-Max-DSpark
A DSpark draft model for speculative decoding with Nex-N2.5-Max.
This repository contains only the draft model. It is not a standalone language model: it is loaded alongside Nex-N2.5-Max and proposes several tokens per step, which the target model then verifies. The output distribution is that of Nex-N2.5-Max; the draft model only reduces latency.
Model Details
| Target model | Nex-N2.5-Max |
| Algorithm | DSpark |
| Parameters | 77.5B |
| Precision | BF16 |
| Draft block size | 5 |
| Modality | Text |
Training
This model was trained with the SpecForge library on a mix of Nex-AGI's in-house data and OpenPerfectBlend.
Deployment
Serve Nex-N2.5-Max with the draft model attached using sglang. Speculative decoding with this draft model requires our customized sglang fork, which is preinstalled in the prebuilt Docker image nexagi/sglang:v0.5.18-nex-patch:
python -m sglang.launch_server \
--model-path <NEX_N2_5_MAX_MODEL_PATH> \
--chat-template <NEX_N2_5_MAX_CHAT_TEMPLATE_PATH> \
--speculative-draft-model-path <NEX_N2_5_MAX_DSPARK_MODEL_PATH> \
--speculative-algorithm DSPARK \
--tp-size <TP_SIZE> \
--mem-fraction-static 0.80 \
--host 0.0.0.0 \
--port 30000
For the target model's own deployment options, sampling parameters, and thinking modes, see the Nex-N2.5-Max model card.