K2-Horizon-375B-A23B
K2-Horizon-375B-A23B is the flagship of the K2-Horizon family: a sparse Mixture-of-Experts model that stores 375B parameters and runs 23B per token, with a 512K context window. We have released the final checkpoint; intermediate checkpoints, along with the data and the training code, will be released.
K2-Horizon-375B-A23B Highlights
- Frontier-class agentic performance. On agentic tool use, terminal, and long-horizon workflow benchmarks it matches or beats open-weight MoE models up to 2.6× its size and is competitive with closed frontier models (see Benchmark Results).
- 512K context. Native 524,288-token context from the midtraining stages onward.
- Intermediate checkpoints. Intermediate checkpoints will be released so capability changes can be studied across training rather than at a single checkpoint.
- Fully open. Training data/recipe and the training code will be made public.
Benchmark Results
Open-weight modelsClosed models
K2-Horizon-375B-A23BNemotron 3 UltraInkling (xhigh)MiniMax-M3GLM 5.2 (max)GPT 5.6 Luna (max)GPT 5.6 Terra (high)Claude Sonnet5 (max)
Params375B550B975B428B753B------
Activated params23B55B41B23B40B------
ArchitectureMoEMoEMoEMoEMoEClosedClosedClosed
Agents
GDPVal-AA
Real-world professional tasks (Elo)
1,4411,1621,2341,3801,4981,5691,5031,584
tau3-Banking
Agentic tool use
34.014.229.115.334.631.128.737.3
Toolathlon Verified
Agentic tool use
65.334.345.553.759.967.564.871.6
Automation Bench Public
Workflow automation
25.38.012.820.526.233.528.034.7
Apex-Agents (pass@1)
Long-horizon professional workflows
24.89.019.023.826.928.625.431.7
MCPMark
MCP tool use
67.745.751.248.872.466.974.065.3
BrowseComp
Deep web research
72.844.477.183.5--83.3--84.7
WildClawBench
In-the-wild agentic tasks
50.934.252.356.455.050.460.0--
Coding
Terminal-Bench 2.1
Agentic terminal use
70.253.955.165.277.980.975.780.5
SciCode
Scientific coding
42.739.946.145.450.552.550.153.6
SWE-Atlas-QnA (strict)
Repo-level code Q&A
48.4--25.542.346.4------
SWE Bench Pro (strict)
Software engineering
42.638.743.143.846.748.8----
Scientific Reasoning
Humanity's Last Exam (without tools)
Expert-level reasoning
32.028.431.939.041.139.538.541.3
GPQA Diamond
Graduate-level science QA
87.386.787.292.989.591.189.691.1
CritPt
Frontier physics reasoning
8.63.15.43.720.921.022.916.9
General
AA-LCR
Long-context reasoning
76.071.073.380.376.778.373.377.0
AA-Omniscience Accuracy
Factual accuracy
23.023.042.017.024.043.045.040.0
AA-Omniscience Non-Hallucination
Non-hallucination rate
74.770.032.082.074.07.010.061.0
Scores in %. SWE-Atlas-QnA and SWE Bench Pro: strict = no internet. BrowseComp: different models use different harness, we use the Discard-all@95k context length proposed in DeepSeek-V3.2 technical report. WildClawBench: we use a subset of the English text-only-modality tasks. Apex-Agents: we use a subset of text-only-modality tasks. GDPVal-AA is the Elo rating.
Quickstart
Serving
vLLM, recipe at recipes.vllm.ai/IFM:
vllm serve IFM/K2-Horizon-375B-A23B \
--revision main \
--tensor-parallel-size 8 \
--enable-expert-parallel \
--trust-remote-code \
--dtype bfloat16 \
--max-model-len 131072 \
--reasoning-parser k2_horizon \
--tool-call-parser k2_horizon \
--enable-auto-tool-choice
SGLang recipe validated on 8× H200 in the SGLang K2 Horizon cookbook:
python3 -m sglang.launch_server \
--model-path IFM/K2-Horizon-375B-A23B \
--revision main \
--tp 8 \
--ep 8 \
--dtype bfloat16 \
--attention-backend fa3 \
--model-loader-extra-config '{"enable_multithread_load":false}' \
--reasoning-parser k2_horizon \
--tool-call-parser k2_horizon \
--host 0.0.0.0 --port 30000
API Usage
[!Tip]
Recommended settings:reasoning_effort="high",temperature=1.0,top_p=0.95.
Reasoning depth is selected per request throughchat_template_kwargs. Thinking is returned inreasoning_contentand the answer incontent.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="IFM/K2-Horizon-375B-A23B",
messages=[{"role": "user", "content": "Explain the result step by step."}],
temperature=1.0,
top_p=0.95,
max_tokens=32768,
extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
message = response.choices[0].message
print("Reasoning:", getattr(message, "reasoning_content", None))
print("Answer:", message.content)
Transformers
Validated with Transformers 4.57.6, PyTorch 2.13.0, Safetensors 0.8.0.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "IFM/K2-Horizon-375B-A23B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id, device_map="auto", dtype="bfloat16", low_cpu_mem_usage=True, trust_remote_code=True
)
inputs = tokenizer("Explain why long-context evaluation is difficult.", return_tensors="pt").to(model.device)
inputs.pop("token_type_ids", None)
outputs = model.generate(**inputs, max_new_tokens=32768, temperature=1.0, top_p=0.95, do_sample=True)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Best Practices
1. Reasoning effort: always high. All reported results use high reasoning effort. Pass {"chat_template_kwargs": {"reasoning_effort": "high"}} on every request.
2. Sampling parameters. temperature=1.0, top_p=0.95.
3. Serving. Use the validated SGLang recipe above: BF16, TP=8 on one 8× H200 node, FlashAttention-3, with multithreaded weight loading disabled. Full recipes for every K2-Horizon size, with measured H200 latency and throughput, are in the SGLang cookbook and the vLLM recipe.
4. Parsers. Enable the k2_horizon reasoning parser for chat, and add the k2_horizon tool-call parser for agent use. Leave both off for plain completion-style generation.
Citation
@misc{k2horizon2026,
title = {Introducing K2 Horizon: Frontier Performance, Radically Open},
author = {{IFM Team}},
year = {2026},
url = {https://ifm.ai/blog/k2/},
}