K2-Horizon-MoVA-36B-A4B-DERISKED-BF16
Behaviorally modified BF16 research checkpoint derived from IFM/K2-Horizon-MoVA-36B-A4B.
Release status
| Item | Status |
|---|---|
| Weight package | 48 BF16 safetensors shards included |
| Structural and full-load smoke validation | Complete |
| Native K2 Horizon configuration, tokenizer, custom code, and chat template | Included |
| Baked deployment system prompt | None |
| Operator qualitative selection | Complete for continued validation |
| Refusal evaluation on the exact release artifact | Pending |
| Capability, long-context, and agent/tool evaluation | Pending |
| Modification recipe | Proprietary and intentionally not distributed |
| Access | Private validation; public manual gating pending results |
DERISKED identifies the Blackfrost research family. It is not a claim of zero refusals, complete safety, harmlessness, unchanged benchmark performance, or production readiness. Quantitative claims will be added only after evaluation of the exact uploaded artifact is complete.
Overview
K2-Horizon-MoVA-36B-A4B-DERISKED-BF16 is an independent Blackfrost research derivative of IFM/K2-Horizon-MoVA-36B-A4B. It retains the upstream K2 Horizon architecture, tokenizer, native chat template, long-context configuration, Mixture-of-Experts feed-forward blocks, and Mixture-of-Values attention layout.
The behavioral modification is encoded in the released weights. It is not a system-prompt wrapper, adapter, runtime filter, or decoding-time intervention. Internal direction data, capture data, evaluation prompts, intermediate checkpoints, target maps, and the reproduction recipe are not included.
Model specifications
| Property | Value |
|---|---|
| Architecture | K2HorizonForCausalLM / k2_horizon |
| Parameters | 36B total / 4B active per token, per the upstream card |
| Transformer layers | 48 |
| Hidden size | 2,560 |
| Dense opening layers | 3 |
| MoE layers | 45 |
| Routed experts | 100 total / 8 active per token |
| Shared experts | 1 per MoE layer |
| MoVA value experts | 64 total / 4 active per token |
| Native maximum context | 524,288 tokens |
| Vocabulary | 250,624 tokens |
| Weight dtype | BF16 |
| Indexed tensor bytes | 74,889,584,040 |
| Weight shards | 48 |
The architectural context limit is not a guarantee that a particular runtime or hardware configuration can allocate or serve the full window.
Lineage
IFM/K2-Horizon-MoVA-36B-A4B@7730b92d1b574e04663b04023d5d6fa83475432f
└── Blackfrost-AI/K2-Horizon-MoVA-36B-A4B-DERISKED-BF16
The derivative was built from the pinned upstream revision shown above. Architecture, tokenizer, chat-template, custom-code, data, and license lineage remain with IFM; the checkpoint contains Blackfrost weight-level behavioral modifications.
Prompting and chat template
The package retains the native K2 Horizon chat template. It supports role-structured messages, reasoning-effort selection, and K2 Horizon tool-call formats. No Blackfrost, Frosty, compliance, or deployment-specific system prompt is baked into this release. A caller may still supply its own system message at inference time.
The upstream recommended generation settings are:
reasoning_effort="high"temperature=1.0top_p=0.95
K2 Horizon supports json, xml, and xml_typed tool-call formatting through chat-template arguments. Production tool use requires a serving runtime with compatible k2_horizon reasoning and tool-call parsers. Applications remain responsible for tool execution, result reinjection, authentication, authorization, conversation state, and output handling.
Transformers quickstart
The checkpoint includes custom model code, so review it and use trust_remote_code=True only in an appropriately isolated environment.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Blackfrost-AI/K2-Horizon-MoVA-36B-A4B-DERISKED-BF16"
tokenizer = AutoTokenizer.from_pretrained(
model_id,
trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
low_cpu_mem_usage=True,
trust_remote_code=True,
)
messages = [{"role": "user", "content": "Explain why long-context evaluation is difficult."}]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
chat_template_kwargs={"reasoning_effort": "high", "tool_call_format": "xml"},
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
inputs.pop("token_type_ids", None)
outputs = model.generate(
**inputs,
max_new_tokens=4096,
temperature=1.0,
top_p=0.95,
do_sample=True,
)
print(tokenizer.decode(outputs[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
The full BF16 checkpoint requires substantial accelerator memory. Sharding, expert parallelism, KV-cache precision, and maximum context must be selected for the target hardware.
Serving notes
The upstream K2 Horizon project documents vLLM and SGLang serving with --trust-remote-code, BF16 weights, expert parallelism, and the k2_horizon reasoning/tool parsers. Treat those upstream recipes as the starting point for this architecture, then validate the exact Blackfrost artifact and runtime version before production use.
Current Blackfrost validation confirms full model loading, finite logits, exact served-model identity, native-template generation, and non-stream reasoning separation. Generated OpenAI message.tool_calls, streaming reasoning/tool deltas, maximum-context allocation, and production concurrency remain pending for the exact release artifact.
Validation status
Completed checks include:
- exact 48-shard indexed inventory and expected logical tensor bytes;
- source config and index provenance;
- complete targeted-write audit during construction;
- readable safetensors and a full BF16 model load;
- finite-logit forward pass;
- native chat-template tokenization at supported reasoning-effort levels;
- served health, exact model identity, and coherent external completion;
- operator qualitative selection for continued evaluation.
Pending checks include refusal behavior, broad capability retention, coding, multilingual behavior, long-context behavior, tool calling, streaming parser behavior, throughput, concurrency, and evaluation of any later quantized derivative. Upstream benchmark results must not be attributed to this modified checkpoint.
Access
This repository is private while exact-artifact validation is in progress. A public release, if approved, is intended to use manual access gating. Access to weights does not imply suitability for a workload or transfer deployment responsibility to Blackfrost.
Limitations and responsibility
- This is an experimental research checkpoint.
- Generated content may be inaccurate, insecure, offensive, or otherwise unsuitable.
- The model is not a security boundary, policy engine, authorization mechanism, or substitute for professional judgment.
- Treat generated text, code, URLs, tool arguments, file paths, and commands as untrusted until independently reviewed.
- Operators are responsible for access controls, monitoring, legal compliance, and safeguards appropriate to their environment.
License and attribution
The upstream repository identifies K2 Horizon as Apache-2.0. Review the upstream model card and included license metadata before use or redistribution.
Blackfrost is independent of and is not affiliated with, sponsored by, or endorsed by IFM. This research artifact is provided as-is, without warranties.
Contact
For reproducible artifact issues, use this repository's Discussions. Do not post credentials, private prompts, personal information, or infrastructure details.