Ornith-1.5-397B-DFlash
Ornith-1.5-397B-DFlash pairs the Ornith-1.5-397B model with a DFlash draft model for speculative decoding. This repository provides the DFlash draft checkpoint and serving configuration for accelerating Ornith-1.5-397B inference. The DFlash checkpoint is not a standalone language model. It is designed to run together with the target model under a DFlash-enabled inference setup.
Ornith-1.5-397B
🐦 Ornith-1.5 is a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For more details on the task, harness, and rollout reward design, please refer to the Ornith-1.5 blog.
DFlash acceleration
DFlash provides the speculative-decoding component of this project. It uses a lightweight block-diffusion draft model to propose multiple tokens in parallel, while Ornith-1.5-397B verifies those proposals as the target model. This allows the serving stack to improve decoding throughput while preserving the target model's output distribution.
In this setup, the two models have distinct roles:
- Ornith-1.5-397B: target model responsible for verification and final generation.
- Ornith-1.5-397B-DFlash: draft model used to generate speculative token blocks for acceleration.
The sections below therefore separate target-model quality benchmarks from DFlash serving configuration and acceleration measurements.
Quickstart
📝 NOTEOrnith-1.5-397B is a reasoning model: by default the assistant turn opens with a <think> … </think> block before the final answer. The serving recipes below enable a reasoning parser so the chain-of-thought is returned in a separate reasoning_content field, and a tool-call parser so the model's <tool_call> blocks are surfaced as OpenAI-style tool_calls.
Serving Ornith-1.5-397B with DFlash requires recent runtimes:
- Transformers ≥ 5.8.1
- vLLM ≥ 0.20.2 (tested with 0.28.0)
- SGLang = 0.5.18
Recommended sampling parameters:
- For general tasks:
temperature=0.6,top_p=0.95,top_k=20 - To reproduce the reported benchmarks:
temperature=1.0
DFlash Installation and Launch
The DFlash draft must be loaded together with the target model. The configurations below use the runtime versions validated for this deployment:
- vLLM:
>=0.20.2(tested with0.28.0) - SGLang:
0.5.18
Replace <DFlash directory> with the local directory or model repository containing this DFlash draft checkpoint.
vLLM with DFlash
vllm serve ornith-ai/Ornith-1.5-397B --served-model-name x --trust-remote-code \
--tensor-parallel-size 1 --port 8801 \
--gpu-memory-utilization 0.85 --max-model-len 32768 \
--speculative-config '{"method":"dflash","model":"<DFlash directory>","num_speculative_tokens":8}'
SGLang with DFlash
python -m sglang.launch_server --model-path ornith-ai/Ornith-1.5-397B --trust-remote-code \
--tp-size 1 --port 30010 --mem-fraction-static 0.80 \
--speculative-algorithm DFLASH \
--speculative-draft-model-path <DFlash directory> \
--speculative-dflash-block-size 8
The examples above use a speculative block size / token count of 8. Adjust tensor parallelism and memory settings to match the actual target-model deployment; the full 397B target model may require multi-GPU serving depending on precision and hardware.
Citation
If you find our work helpful, feel free to give us a cite.
@misc{ornith_1_5,
title = {{Ornith-1.5}: From Self-Scaffolding to Self-Improvement},
url = {https://ornith.ai/ornith_1_5.html},
author = {{Ornith Team}},
year = {2026}
}