Decision-1.0-Kai-0.6B
Kai, from kairos — the right moment to choose.
Open Decision Foundation Models
Choose an action, judge a condition, or score against your own rubric. Kai reads the context and candidate descriptions together, then returns structured decisions and probability distributions.
Decision collection · Download
Measured decisions
53.52 overall — above both general Laya models, and +7.04 points over the previous Kai release.
| Benchmark group | Kai · 0.6B | Laya English | Laya Multilingual |
|---|---|---|---|
| Decisions | 57.96 | 56.54 | 47.25 |
| Composition | 40.83 | 35.33 | 38.92 |
| Reading | 54.69 | 51.41 | 50.78 |
| Inference | 69.79 | 63.75 | 57.29 |
| Transfer | 48.37 | 53.06 | 47.13 |
| Weighted overall | 53.52 | 51.03 | 47.19 |
On the benchmark's 160-question BoolQ slice, Kai scores 74.38%, versus 69.38% for each Laya reference.
Same 54 tasks and 3,766 scored questions, with fixed source/family weights. Both Laya references use released general weights, without target-dataset fine-tuning. This is an observed regression suite; results vary by task. Full matrix, paired intervals, and methods.
Three ways to decide
| Choice | Noul | Score |
|---|---|---|
| Choose among your actions or categories. | Check whether a condition is true. | Rate against ordered criteria. |
| Candidate IDs and probabilities. | Probability of yes. | Level distribution and expected score. |
Use the System One format: state / model / questions → answers. Ask many questions about one context, or apply shared questions across a batch of contexts. Candidates are supplied at runtime.
SDK and curl examples · Training-source attribution
128 mixed questions in 163 ms — 58% lower latency. Automatic typed scheduling accelerates the measured SystemOne runtime with the same weights. Paired local AMD measurements on a fixed workload. Latency and scaling.
Download for local inference
hf download vllm-sr/Decision-1.0-Kai-0.6B --local-dir Decision-1.0-Kai-0.6B
This repository contains model files, provenance and the Transformers loading code. Serving requires a compatible vLLM Semantic Router Decision runtime, distributed separately. Check its hardware support before serving. The runtime must support vllm-sr-decision format version 1 and the file map in config.json. For local inference without a server, see Use with 🤗 Transformers below.
Use with 🤗 Transformers
The repository includes its inference code, so stock Transformers can download and run the complete model locally with trust_remote_code=True. system_one takes and returns the same System One request and response bodies as the Decision runtime; nothing is generated.
pip install "transformers>=4.57" torch safetensors huggingface_hub
from transformers import AutoModel
model = AutoModel.from_pretrained("vllm-sr/Decision-1.0-Kai-0.6B", trust_remote_code=True)
response = model.system_one(
state="The parcel arrived damaged. Please send a replacement today.",
questions={
"route": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {"delivery": "Damaged or missing parcels", "billing": "Payments and invoices"},
},
"urgent": {
"type": "noul",
"instructions": "Does the customer request action today?",
},
},
)
print(response["answers"]["route"]["choice"], response["answers"]["urgent"]["noul"])
pipeline("decision", model="vllm-sr/Decision-1.0-Kai-0.6B", trust_remote_code=True) accepts the same request body. A malformed question is answered with an invalid_question error. The model loads on the first GPU when one is visible, otherwise on the CPU (pass device="cpu" or device="cuda:0" to choose); weights and arithmetic are FP32. A complete question, its candidates and the state are limited to 1,024 tokens; if a question is longer, every question of the request is answered with a max_length_exceeded error and nothing is truncated.
Use
Replace the placeholder with a SystemOne-compatible endpoint configured to serve Decision-1.0-Kai-0.6B, and set DECISION_API_KEY to that endpoint's key.
pip install typesafe-sdk
import os
from typesafe_sdk import TypeSafeClient, Choice, Noul
client = TypeSafeClient(
api_key=os.environ["DECISION_API_KEY"],
base_url="https://your-decision-endpoint.example",
model="Decision-1.0-Kai-0.6B",
)
questions = {
"route": Choice(instructions="Which team should handle this request?",
criteria={"delivery": "Damaged or missing parcels", "billing": "Payments and invoices"}),
"urgent": Noul(instructions="Does the customer request action today?"),
}
response = client.system_one(state="The parcel arrived damaged. Please send a replacement today.", questions=questions)
print(response.choices["route"].choice, response.nouls["urgent"].noul)
The same request with curl:
curl -X POST https://your-decision-endpoint.example/v1/systemone \
-H "Authorization: Bearer $DECISION_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"model": "Decision-1.0-Kai-0.6B",
"state": "The parcel arrived damaged. Please send a replacement today.",
"questions": {
"route": {"type": "choice", "instructions": "Which team should handle this request?", "criteria": {"delivery": "Damaged or missing parcels", "billing": "Payments and invoices"}},
"urgent": {"type": "noul", "instructions": "Does the customer request action today?"}
}
}'
Official Python SDK · HTTP API
Architecture
Three 22-layer bidirectional paths share multilingual input embeddings. Each decision type has its own interaction layers and candidate readout. Candidates within a question are scored together; questions are processed in batches.
Choice and Score were updated while preserving the released Noul path exactly. The model files retain the three-path architecture and complete-input contract.
Architecture and readout diagrams · Training and release notes
The complete 1,024-token budget includes context, questions, candidates, and special tokens. Longer inputs are rejected. This release's headline evaluation includes English and Chinese; broader multilingual results from previous weights are historical evidence.
Transfer remains behind Laya English; Reading is 1.25 points below the previous Kai release. Probabilities are not guarantees, and candidate order can affect predictions. Full results and limitations · AMD runtime measurements
Built on Vela Encoder. Attribution · License scope