K2-Horizon-MoVA-36B-A4B-DERISKED-NVFP4
Experimental NVFP4 deployment derivative of Blackfrost-AI/K2-Horizon-MoVA-36B-A4B-DERISKED-BF16.
Experimental public build: conversion, exhaustive structural checks, full runtime load, coherent generation, reasoning separation, streaming, multi-turn history, and repeated native tool-call parsing have passed on the exact uploaded artifact. Behavioral and full-window evaluation remain in progress.
Release status
| Item | Status |
|---|---|
| Parent BF16 checkpoint | Public |
| NVFP4 conversion | Passed: ModelOpt NVFP4 W4A4 with BF16 MoVA value experts |
| Structural verification | Passed: index/shards, tensor inventory, required dense tensors, and exhaustive finite-value scan |
| Full-load smoke | Passed on NVIDIA RTX PRO 6000 Blackwell Server Edition |
| Coherent generation | Passed |
| Reasoning separation | Passed |
| Streaming reasoning/content separation | Passed |
| Native tool-call parsing | Passed: offline fixtures and three consecutive live OpenAI API trials |
| Refusal evaluation | Pending |
| Capability-retention evaluation | Pending |
| GB10 single-request throughput | Passed: 21.38 wall-clock completion tok/s median across three runs |
| Full-window prompt validation | Pending; 262,144-token allocation passed |
| Baked deployment system prompt | None planned |
| Modification/conversion recipe | Proprietary and intentionally not distributed |
| Access | Public experimental artifact |
DERISKED identifies the Blackfrost research family. It is not a claim of zero refusals, complete safety, harmlessness, lossless quantization, benchmark parity, or production readiness.
Overview
This repository contains a Blackfrost NVFP4 deployment derivative produced from Blackfrost-AI/K2-Horizon-MoVA-36B-A4B-DERISKED-BF16. The parent checkpoint is itself a weight-level behavioral derivative of IFM/K2-Horizon-MoVA-36B-A4B.
The package retains the native K2 Horizon tokenizer, chat template, long-context configuration, custom architecture code, MoE layout, and MoVA layout while representing eligible weights in a ModelOpt-compatible NVFP4 deployment format. K2 Horizon's self_attn.v_experts matrices remain BF16 because the native fused-MoVA runtime consumes dense value-expert matrices; routers, embeddings, normalization tensors, and the LM head also retain their runtime-appropriate higher precision.
Internal behavioral directions, capture data, target maps, calibration material, intermediate checkpoints, evaluation prompts, and reproduction recipes will not be distributed.
Model specifications
| Property | Value |
|---|---|
| Parent architecture | K2HorizonForCausalLM / k2_horizon |
| Parameters | 36B total / 4B active per token, per the upstream card |
| Transformer layers | 48 |
| Hidden size | 2,560 |
| Routed experts | 100 total / 8 active per token |
| Shared experts | 1 per MoE layer |
| MoVA value experts | 64 total / 4 active per token |
| Native maximum context | 524,288 tokens |
| Deployment format | Mixed ModelOpt NVFP4 W4A4 / BF16 MoVA value experts, with FP8 KV-cache metadata |
| ModelOpt source | 0.48.0dev tag at 022767c7ab3d7d36211affd85e5c496770cde768 |
| Weight shards | 4 |
| Safetensors bytes | 33,794,653,520 |
| Tensor inventory | 58,515 total: 3,159 BF16; 27,678 FP32; 13,839 FP8 E4M3FN; 13,839 UINT8 |
| Dense MoVA value-expert weights | 2,880 BF16 tensors |
| Conversion smoke runtime | vLLM 0.27.1+aeon.sm120.rtx; PyTorch 2.13.0+cu130; Transformers 5.14.1 |
| Spark deployment runtime | vLLM 0.1.dev20073+g8e685d198; PyTorch 2.13.0+cu130; Transformers 5.15.1 |
| Validated hardware | NVIDIA RTX PRO 6000 Blackwell Server Edition, 96 GB; NVIDIA GB10 Spark |
| Spark served window | 262,144 tokens configured and allocated; full-window prompt validation pending |
The native context limit is not a guarantee that the eventual runtime or hardware configuration can allocate or serve the full window.
Lineage
IFM/K2-Horizon-MoVA-36B-A4B@7730b92d1b574e04663b04023d5d6fa83475432f
└── Blackfrost-AI/K2-Horizon-MoVA-36B-A4B-DERISKED-BF16
└── Blackfrost-AI/K2-Horizon-MoVA-36B-A4B-DERISKED-NVFP4
Prompting and runtime compatibility
The artifact retains the native K2 Horizon chat template, with reasoning_effort="high", temperature=1.0, and top_p=0.95 as the upstream starting settings. No Blackfrost, Frosty, compliance, or deployment-specific system prompt is baked into the package.
The exact artifact passed native K2 Horizon reasoning separation, streaming reasoning/content separation, generic OpenAI multi-turn history, and repeated schema-valid tool-call parsing through an OpenAI-compatible vLLM endpoint. The deployment compatibility template defaults tool calls to K2's native JSON representation; clients may still request K2 XML or typed XML explicitly. Generic NVFP4 support alone does not imply support for K2 Horizon custom code, expert routing, MoVA, parsers, streaming, or long-context allocation.
GB10 OpenAI API deployment kit
The repository includes a self-contained NVIDIA GB10 vLLM deployment kit with:
- the exact prepacked-MoVA model adapter used for the accepted run;
- K2 reasoning and tool-call parsers;
- the OpenAI multi-turn compatibility chat template;
- an immutable validated ARM64 container reference;
- a portable
launch.shwith no machine-specific paths; - a standard-library API verifier for reasoning and tool calls.
The patch removes a decode-time reconstruction of all 64 BF16 MoVA value experts in each of 45 sparse layers. On the accepted GB10 test, median wall-clock completion throughput improved from 6.06 to 21.38 tokens/s, or 3.53x. These rates include prefill, time to first token, detokenization, and local HTTP transport.
Validation record and remaining gates
Completed on the exact artifact:
- converter, ModelOpt, CUDA, framework, driver, runtime, and hardware recorded;
- exact quantization map, four-shard inventory, byte count, dtype inventory, and exhaustive finite-value scan;
- required dense MoVA tensors and native tokenizer/template/custom-code assets preserved;
- full native runtime load and coherent generation;
- exact single served-model identity and separated reasoning output;
- streaming reasoning/content separation with clean termination;
- native tool-call parser fixtures and three consecutive live calls with valid JSON arguments;
- three-run GB10 throughput benchmark at 21.38 wall-clock completion tok/s median.
Still pending:
- refusal and red-line release gates;
- capability-retention comparison against the exact parent BF16 checkpoint;
- full-window prompt validation and concurrency saturation under representative traffic.
No upstream or BF16-parent benchmark result may be attributed to this quantized derivative without rerunning the exact artifact.
Access
This NVFP4 repository is a public experimental artifact. Public availability does not imply production readiness or completion of the pending behavioral and full-window gates.
Limitations and responsibility
- This is an experimental quantized research checkpoint.
- Quantization can alter capability, numerical behavior, refusal behavior, formatting, routing, and long-context stability.
- Generated content may be inaccurate, insecure, offensive, or otherwise unsuitable.
- The model is not a security boundary, policy engine, authorization mechanism, or substitute for professional judgment.
- Operators are responsible for access controls, monitoring, legal compliance, and safeguards appropriate to their environment.
License and attribution
The upstream K2 Horizon repository identifies the model license as Apache-2.0. Review the upstream model card and the BF16 parent repository before use or redistribution.
Blackfrost is independent of and is not affiliated with, sponsored by, or endorsed by IFM. This artifact is provided as-is, without warranties.
Contact
For reproducible artifact issues, use this repository's Discussions. Do not post credentials, private prompts, personal information, or infrastructure details.