nex-agi/Nex-N2.5-Max-DSpark

🤗 Hugging Face 来源text-generationapache-2.077.5B 参数155 GBsafetensors✓ 74 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo nex-agi/Nex-N2.5-Max-DSpark ./model-folder
需要做种者 →

💻 GitHub  ·   🤗 Hugging Face  ·   🌐 Website

Nex-N2.5-Max-DSpark

A DSpark draft model for speculative decoding with Nex-N2.5-Max.

This repository contains only the draft model. It is not a standalone language model: it is loaded alongside Nex-N2.5-Max and proposes several tokens per step, which the target model then verifies. The output distribution is that of Nex-N2.5-Max; the draft model only reduces latency.

Model Details

Target model Nex-N2.5-Max
Algorithm DSpark
Parameters 77.5B
Precision BF16
Draft block size 5
Modality Text

Training

This model was trained with the SpecForge library on a mix of Nex-AGI's in-house data and OpenPerfectBlend.

Deployment

Serve Nex-N2.5-Max with the draft model attached using sglang. Speculative decoding with this draft model requires our customized sglang fork, which is preinstalled in the prebuilt Docker image nexagi/sglang:v0.5.18-nex-patch:

python -m sglang.launch_server \
  --model-path <NEX_N2_5_MAX_MODEL_PATH> \
  --chat-template <NEX_N2_5_MAX_CHAT_TEMPLATE_PATH> \
  --speculative-draft-model-path <NEX_N2_5_MAX_DSPARK_MODEL_PATH> \
  --speculative-algorithm DSPARK \
  --tp-size <TP_SIZE> \
  --mem-fraction-static 0.80 \
  --host 0.0.0.0 \
  --port 30000

For the target model's own deployment options, sampling parameters, and thinking modes, see the Nex-N2.5-Max model card.