FrontiersMind/Lumma-fev-9b

🤗 Hugging Face 来源text-classificationapache-2.07.9B 参数16 GBsafetensors✓ 2 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo FrontiersMind/Lumma-fev-9b ./model-folder
需要做种者 →

Lumma-Fev-9B

Lumma-Fev-9B is a decision model. It reads one document (the state) and a set of typed questions, and returns a probability distribution for each question in a single forward pass. It never generates text, so there is nothing to parse and nothing to hallucinate.

Benchmarks

Benchmark / Metric TypeSafe Jev 1.13.0 Laya Gliner2.5-Decide Lumma-Fev-0.15B Lumma-Fev-0.6B Lumma-Fev-4B Lumma-Fev-9B
Typed-decisions 0.72 0.76 0.46 0.49 0.64 0.78 0.81
AG News 0.91 0.95 0.48 0.89 0.84 0.93 0.95
DAIR Emotion 0.48 0.595 0.55 0.68 0.89 0.94 0.94
Banking77 0.87 0.425 0.68 0.47 0.90 0.93 0.94
Average 0.75 0.68 0.54 0.63 0.82 0.90 0.91
P50 latency 256 ms 32.8 ms 99.85 ms 35.576 ms 45.82 ms 175.93 ms 269.39 ms
Weights Close Open Open Open Open Open Open
Cost $0.042 / 1M tokens $0 self-hosted $0 self-hosted $0 self-hosted $0 self-hosted $0 self-hosted $0 self-hosted

Use it with transformers

from transformers import AutoModel
import torch

model = AutoModel.from_pretrained("FrontiersMind/Lumma-fev-9b", trust_remote_code=True,dtype=torch.bfloat16)   # the tokenizer is loaded with it

answers = model.decide(
    state="My running shoes arrived in the wrong size. Can I exchange them for a size 10?",
    questions={
    "department": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
            "returns": "Exchanges, wrong items, damaged items",
            "shipping": "Delivery status, delays, lost packages",
            "billing": "Charges, invoices, payment problems",
        },
    }
}
)
print(answers)

state can be text, a JSON object or an array. Pin a version with revision="<commit>", and move the model to a GPU with model.to("cuda") (bf16 is the stored precision).

Use it with the lumma-fev package

pip install lumma-fev              # local inference
pip install "lumma-fev[serve]"     # plus the API server
import lumma_fev

model = lumma_fev.load("FrontiersMind/Lumma-fev-9b")        # picks cuda, mps or cpu
print(model.decide("Two charges on my card for one order.",
                   {"billing": {"type": "noul", "instructions": "Is this about billing?"}}))

Serve it as an API

lumma-fev-serve exposes the TypeSafe POST /v1/systemone contract, so existing TypeSafe clients work by changing their base URL.

lumma-fev-serve --model FrontiersMind/Lumma-fev-9b --host 0.0.0.0 --port 8000
# optional: LUMMA_FEV_API_KEY=<key> requires "Authorization: Bearer <key>"; --cors for browser apps
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
  "state": "I was charged twice. Please fix this ASAP.",
  "questions": {"billing": {"type": "noul", "instructions": "Is this ticket about billing?"}}
}'

From Python, with the lumma-fev client or the TypeSafe SDK:

from lumma_fev import Client

client = Client("http://127.0.0.1:8000")
print(client.decide("I was charged twice.", {"billing": {"type": "noul", "instructions": "Is this about billing?"}}))

from typesafe_sdk import Noul, TypeSafeClient     # pip install typesafe-sdk

with TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8000") as ts:
    print(ts.system_one(state="I was charged twice.", questions={"billing": Noul(instructions="Is this about billing?")}).nouls["billing"].noul)

Answers

Type Criteria Answer
noul optional {"true": ..., "false": ...} noul: probability of yes
choice {name: description or null} choice (most likely name), confidence, probabilities by name
score ordered list of levels score (expected level), confidence, legend, probabilities by level

Questions never see each other: each one reads the state and its own instructions and options only, so one question cannot change another's answer. Text inside the request cannot forge the model's delimiter tokens.

Limits

  • confidence and probabilities are the model's own estimates. Measure calibration on your own labelled data before gating automated actions on them.

NOTE: This Model is continual pre-trained and fine-tuned on top Qwen-3.5.-4B Series.

Feedback