ornith-ai/Ornith-1.5-9B-DFlash

🤗 Hugging Face 来源mit1.3B 参数2.6 GBsafetensors✓ 3 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo ornith-ai/Ornith-1.5-9B-DFlash ./model-folder
需要做种者 →

Ornith-1.5-9B-DFlash

Ornith-1.5-9B-DFlash pairs the Ornith-1.5-9B model with a DFlash draft model for speculative decoding. This repository provides the DFlash draft checkpoint and serving configuration for accelerating Ornith-1.5-9B inference. The DFlash checkpoint is not a standalone language model. It is designed to run together with the target model under a DFlash-enabled inference setup.

Ornith-1.5-9B

🐦 Ornith-1.5 is a major step toward building foundation models through end-to-end self-improvement.
Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For more details on the task, harness, and rollout reward design, please refer to the Ornith-1.5 blog.

Ornith-1.5-9B is the most lightweight member of the Ornith-1.5 family — a dense 9B model designed for efficient single-GPU deployment, and edge-deployable on mobile devices via its quantized Ornith-1.5-9B-Mobile variant.

DFlash acceleration

DFlash provides the speculative-decoding component of this project. It uses a lightweight block-diffusion draft model to propose multiple tokens in parallel, while Ornith-1.5-9B verifies those proposals as the target model. This allows the serving stack to improve decoding throughput while preserving the target model's output distribution.

In this setup, the two models have distinct roles:

  • Ornith-1.5-9B: target model responsible for verification and final generation.
  • Ornith-1.5-9B-DFlash: draft model used to generate speculative token blocks for acceleration.

The sections below therefore separate target-model quality benchmarks from DFlash serving configuration and acceleration measurements.

Quickstart

📝 NOTE

Ornith-1.5-9B is a reasoning model: by default the assistant turn opens with a <think> … </think> block before the final answer. The serving recipes below enable a reasoning parser so the chain-of-thought is returned in a separate reasoning_content field, and a tool-call parser so the model's <tool_call> blocks are surfaced as OpenAI-style tool_calls.

Serving Ornith-1.5-9B with DFlash requires recent runtimes:

  • Transformers ≥ 5.8.1
  • vLLM ≥ 0.20.2 (tested with 0.28.0)
  • SGLang = 0.5.18

Recommended sampling parameters:

  • For general tasks: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
  • For precise coding tasks: temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0

DFlash Installation and Launch

The DFlash draft must be loaded together with the target model. The configurations below use the runtime versions validated for this deployment:

  • vLLM: >=0.20.2 (tested with 0.28.0)
  • SGLang: 0.5.18

Replace <DFlash directory> with the local directory or model repository containing this DFlash draft checkpoint.

vLLM with DFlash

vllm serve ornith-ai/Ornith-1.5-9B --served-model-name x --trust-remote-code \
    --tensor-parallel-size 1 --port 8801 \
    --gpu-memory-utilization 0.85 --max-model-len 32768 \
    --speculative-config '{"method":"dflash","model":"<DFlash directory>","num_speculative_tokens":8}'

SGLang with DFlash

python -m sglang.launch_server --model-path ornith-ai/Ornith-1.5-9B --trust-remote-code \
    --tp-size 1 --port 30010 --mem-fraction-static 0.80 \
    --speculative-algorithm DFLASH \
    --speculative-draft-model-path <DFlash directory> \
    --speculative-dflash-block-size 8

The examples above use a speculative block size / token count of 8. Ornith-1.5-9B is a dense ~9B model (≈19 GB in bf16) designed for single-GPU deployment, so the examples use tensor parallelism of 1. Adjust memory settings as needed for the target model, DFlash draft model, and runtime overhead.

Citation

If you find our work helpful, feel free to give us a cite.

@misc{ornith_1_5,
    title = {{Ornith-1.5}: From Self-Scaffolding to Self-Improvement},
    url = {https://ornith.ai/ornith_1_5.html},
    author = {{Ornith Team}},
    year = {2026}
}