Decision-1.0-Eos-0.8B
Eos — the dawn of a clearer decision.
Open Decision Foundation Models
Turn context into choices, judgments, and scores. Eos reads your questions and candidate descriptions, then returns structured answers with probability distributions. Ask many questions about one context, or apply the same questions across a batch of contexts.
Decision collection · Architecture · Evaluation
Measured decisions
61.89 overall — ahead of all four reference models on the same 54-task benchmark.
| Model | Overall ↑ | Decisions | Composition | Reading | Inference | Transfer |
|---|---|---|---|---|---|---|
| Eos · 0.8B | 61.89 | 65.94 | 46.04 | 70.31 | 81.67 | 52.01 |
| Kev · 0.8B | 58.28 | 60.14 | 42.29 | 67.81 | 68.75 | 61.19 |
| Qwen3.5 · 2B¹ | 57.24 | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 |
| Laya English | 51.03 | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 |
| Laya Multilingual | 47.19 | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 |
3,766 scored questions; fixed task weights and the same requests for every model. Overall weights are 30% / 25% / 15% / 15% / 15%. Both Laya references use their released general weights. ¹Qwen uses the benchmark's fixed decision-letter readout. This is an observed regression benchmark; overall leadership does not imply a win on every task. Full 54-task matrix, intervals, and method.
Your questions define the task
| Decision | What you receive |
|---|---|
| Choice | A distribution over your named actions or categories. |
| Noul | The probability that a condition is true. |
| Score | A distribution over ordered levels and the expected score. |
Use the System One request format: state / model / questions → answers. Candidate descriptions are supplied at runtime. Multiple questions can share a context, and the same questions can be applied to several contexts.
Download and run
Download the complete repository so the backbone, decision head, tokenizer, and configuration files stay together:
hf download vllm-sr/Decision-1.0-Eos-0.8B --local-dir Decision-1.0-Eos-0.8B
Serve this checkpoint with the vLLM Semantic Router Decision runtime. This repository contains model artifacts, documentation and Transformers loading code, with no bundled service. The weights do not depend on a particular accelerator; which hardware can serve them is decided by the runtime. The root config.json describes the Decision artifact layout and names the Transformers loading code. The calibrated temperature is recorded in config.json.
Once a compatible endpoint is serving Eos, make a decision with:
DECISION_BASE_URL=http://127.0.0.1:8000
curl -X POST "$DECISION_BASE_URL/v1/systemone" \
-H 'Content-Type: application/json' \
--data '{
"model": "Decision-1.0-Eos-0.8B",
"state": "The parcel arrived damaged. Please send a replacement today.",
"questions": {
"route": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"delivery": "Damaged or missing parcels",
"billing": "Payments and invoices"
}
}
}
}'
Use with 🤗 Transformers
The repository includes its inference code, so stock Transformers can download and run the complete model locally with trust_remote_code=True. system_one takes and returns the same System One request and response bodies as the Decision runtime; nothing is generated.
pip install "transformers>=5.17" torch safetensors huggingface_hub
from transformers import AutoModel
model = AutoModel.from_pretrained("vllm-sr/Decision-1.0-Eos-0.8B", trust_remote_code=True)
response = model.system_one(
state="Customer requests a refund.",
questions={
"route": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {"billing": "Payments and refunds", "technical": "Product faults"},
},
"priority": {
"type": "score",
"instructions": "How urgent is this request?",
"criteria": ["Not urgent", "Soon", "Today", "Immediately"],
},
},
)
print(response["answers"]["route"]["choice"], response["answers"]["priority"]["score"])
pipeline("decision", model="vllm-sr/Decision-1.0-Eos-0.8B", trust_remote_code=True) accepts the same request body. A malformed question is answered with an invalid_question error. The model loads on the first GPU when one is visible, otherwise on the CPU (pass device="cpu" or device="cuda:0" to choose). On a GPU the backbone runs in BF16 with an FP32 decision head; on the CPU everything runs in FP32. A complete question, its candidates and the state are limited to 16,384 tokens; if a question is longer, every question of the request is answered with a max_length_exceeded error and nothing is truncated. flash-linear-attention speeds up the linear-attention layers on a GPU.
Built to decide
A 24-layer hybrid decoder combines gated linear attention and full attention. A shared decision head reads candidate endpoints against a global query representation and scores all candidates in one forward pass per question batch.
The released inference model contains the text backbone and decision head. See the architecture and readout diagrams for the computation graph.
Built on Qwen3.5-0.8B. English and Chinese are represented in the release evaluation. Transfer, candidate-carried evidence, and some rule tasks remain areas for improvement; probabilities can be overconfident. The maximum complete input is 16,384 tokenizer tokens; this capacity does not establish long-context reasoning quality.