Blackfrost-AI/K2-Horizon-MoVA-36B-A4B-DERISKED-BF16

Verified creator Blackfrost-AI verified
🤗 Hugging Face sourcetext-generationapache-2.037.4B params4B activated75 GBsafetensors✓ 50 checksumsupdated today
Needs seeder →

K2-Horizon-MoVA-36B-A4B-DERISKED-BF16

Behaviorally modified BF16 research checkpoint derived from IFM/K2-Horizon-MoVA-36B-A4B.

Release status

Item Status
Weight package 48 BF16 safetensors shards included
Structural and full-load smoke validation Complete
Native K2 Horizon configuration, tokenizer, custom code, and chat template Included
Baked deployment system prompt None
Operator qualitative selection Complete for continued validation
Refusal evaluation on the exact release artifact Pending
Capability, long-context, and agent/tool evaluation Pending
Modification recipe Proprietary and intentionally not distributed
Access Private validation; public manual gating pending results

DERISKED identifies the Blackfrost research family. It is not a claim of zero refusals, complete safety, harmlessness, unchanged benchmark performance, or production readiness. Quantitative claims will be added only after evaluation of the exact uploaded artifact is complete.

Overview

K2-Horizon-MoVA-36B-A4B-DERISKED-BF16 is an independent Blackfrost research derivative of IFM/K2-Horizon-MoVA-36B-A4B. It retains the upstream K2 Horizon architecture, tokenizer, native chat template, long-context configuration, Mixture-of-Experts feed-forward blocks, and Mixture-of-Values attention layout.

The behavioral modification is encoded in the released weights. It is not a system-prompt wrapper, adapter, runtime filter, or decoding-time intervention. Internal direction data, capture data, evaluation prompts, intermediate checkpoints, target maps, and the reproduction recipe are not included.

Model specifications

Property Value
Architecture K2HorizonForCausalLM / k2_horizon
Parameters 36B total / 4B active per token, per the upstream card
Transformer layers 48
Hidden size 2,560
Dense opening layers 3
MoE layers 45
Routed experts 100 total / 8 active per token
Shared experts 1 per MoE layer
MoVA value experts 64 total / 4 active per token
Native maximum context 524,288 tokens
Vocabulary 250,624 tokens
Weight dtype BF16
Indexed tensor bytes 74,889,584,040
Weight shards 48

The architectural context limit is not a guarantee that a particular runtime or hardware configuration can allocate or serve the full window.

Lineage

IFM/K2-Horizon-MoVA-36B-A4B@7730b92d1b574e04663b04023d5d6fa83475432f
└── Blackfrost-AI/K2-Horizon-MoVA-36B-A4B-DERISKED-BF16

The derivative was built from the pinned upstream revision shown above. Architecture, tokenizer, chat-template, custom-code, data, and license lineage remain with IFM; the checkpoint contains Blackfrost weight-level behavioral modifications.

Prompting and chat template

The package retains the native K2 Horizon chat template. It supports role-structured messages, reasoning-effort selection, and K2 Horizon tool-call formats. No Blackfrost, Frosty, compliance, or deployment-specific system prompt is baked into this release. A caller may still supply its own system message at inference time.

The upstream recommended generation settings are:

  • reasoning_effort="high"
  • temperature=1.0
  • top_p=0.95

K2 Horizon supports json, xml, and xml_typed tool-call formatting through chat-template arguments. Production tool use requires a serving runtime with compatible k2_horizon reasoning and tool-call parsers. Applications remain responsible for tool execution, result reinjection, authentication, authorization, conversation state, and output handling.

Transformers quickstart

The checkpoint includes custom model code, so review it and use trust_remote_code=True only in an appropriately isolated environment.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Blackfrost-AI/K2-Horizon-MoVA-36B-A4B-DERISKED-BF16"

tokenizer = AutoTokenizer.from_pretrained(
    model_id,
    trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    low_cpu_mem_usage=True,
    trust_remote_code=True,
)

messages = [{"role": "user", "content": "Explain why long-context evaluation is difficult."}]
prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    chat_template_kwargs={"reasoning_effort": "high", "tool_call_format": "xml"},
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
inputs.pop("token_type_ids", None)
outputs = model.generate(
    **inputs,
    max_new_tokens=4096,
    temperature=1.0,
    top_p=0.95,
    do_sample=True,
)
print(tokenizer.decode(outputs[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

The full BF16 checkpoint requires substantial accelerator memory. Sharding, expert parallelism, KV-cache precision, and maximum context must be selected for the target hardware.

Serving notes

The upstream K2 Horizon project documents vLLM and SGLang serving with --trust-remote-code, BF16 weights, expert parallelism, and the k2_horizon reasoning/tool parsers. Treat those upstream recipes as the starting point for this architecture, then validate the exact Blackfrost artifact and runtime version before production use.

Current Blackfrost validation confirms full model loading, finite logits, exact served-model identity, native-template generation, and non-stream reasoning separation. Generated OpenAI message.tool_calls, streaming reasoning/tool deltas, maximum-context allocation, and production concurrency remain pending for the exact release artifact.

Validation status

Completed checks include:

  • exact 48-shard indexed inventory and expected logical tensor bytes;
  • source config and index provenance;
  • complete targeted-write audit during construction;
  • readable safetensors and a full BF16 model load;
  • finite-logit forward pass;
  • native chat-template tokenization at supported reasoning-effort levels;
  • served health, exact model identity, and coherent external completion;
  • operator qualitative selection for continued evaluation.

Pending checks include refusal behavior, broad capability retention, coding, multilingual behavior, long-context behavior, tool calling, streaming parser behavior, throughput, concurrency, and evaluation of any later quantized derivative. Upstream benchmark results must not be attributed to this modified checkpoint.

Access

This repository is private while exact-artifact validation is in progress. A public release, if approved, is intended to use manual access gating. Access to weights does not imply suitability for a workload or transfer deployment responsibility to Blackfrost.

Limitations and responsibility

  • This is an experimental research checkpoint.
  • Generated content may be inaccurate, insecure, offensive, or otherwise unsuitable.
  • The model is not a security boundary, policy engine, authorization mechanism, or substitute for professional judgment.
  • Treat generated text, code, URLs, tool arguments, file paths, and commands as untrusted until independently reviewed.
  • Operators are responsible for access controls, monitoring, legal compliance, and safeguards appropriate to their environment.

License and attribution

The upstream repository identifies K2 Horizon as Apache-2.0. Review the upstream model card and included license metadata before use or redistribution.

Blackfrost is independent of and is not affiliated with, sponsored by, or endorsed by IFM. This research artifact is provided as-is, without warranties.

Contact

For reproducible artifact issues, use this repository's Discussions. Do not post credentials, private prompts, personal information, or infrastructure details.