💻 GitHub · 🤗 Hugging Face · 🌐 Website
Nex-N2.5-Pro-DFlash
A DFlash draft model for speculative decoding with Nex-N2.5-Pro, supporting text and multimodal workloads.
This repository contains only the draft model. It is not a standalone language model: it is loaded alongside Nex-N2.5-Pro and proposes a block of tokens per step, which the target model then verifies. The output distribution is that of Nex-N2.5-Pro; the draft model only reduces latency.
Model Details
| Target model | Nex-N2.5-Pro |
| Algorithm | DFlash |
| Parameters | 1.29B |
| Precision | BF16 |
| Draft block size | 16 |
| Modality | Text and multimodal (via the target model) |
Training
This model was trained with the SpecForge library on Nex-AGI's in-house data, together with a portion of OpenPerfectBlend.
Training uses a 0.1 CE + 0.9 TV loss, temperature 0.7, a DFlash block size of 16, and 512 anchors per sample.
Deployment
Serve Nex-N2.5-Pro with the draft model attached using sglang. Speculative decoding with this draft model requires our customized sglang fork, which is preinstalled in the prebuilt Docker image nexagi/sglang:v0.5.18-nex-patch:
python -m sglang.launch_server \
--model-path <NEX_N2_5_PRO_MODEL_PATH> \
--chat-template <NEX_N2_5_PRO_CHAT_TEMPLATE_PATH> \
--speculative-draft-model-path <NEX_N2_5_PRO_DFLASH_MODEL_PATH> \
--speculative-algorithm DFLASH \
--speculative-dflash-block-size 16 \
--tp-size 8 \
--context-length 262144 \
--max-total-tokens 262144 \
--mem-fraction-static 0.80 \
--chunked-prefill-size 8192 \
--host 0.0.0.0 \
--port 30000
For the target model's own deployment options, sampling parameters, and thinking modes, see the Nex-N2.5-Pro model card.