Blackfrost-AI/K2-Horizon-MoVA-36B-A4B-DERISKED-NVFP4

认证创作者 Blackfrost-AI 已认证
🤗 Hugging Face 来源text-generationapache-2.023.1B 参数激活 4B32 GBsafetensors✓ 6 个校验和今天更新
需要做种者 →

K2-Horizon-MoVA-36B-A4B-DERISKED-NVFP4

Experimental NVFP4 deployment derivative of Blackfrost-AI/K2-Horizon-MoVA-36B-A4B-DERISKED-BF16.

Experimental public build: conversion, exhaustive structural checks, full runtime load, coherent generation, reasoning separation, streaming, multi-turn history, and repeated native tool-call parsing have passed on the exact uploaded artifact. Behavioral and full-window evaluation remain in progress.

Release status

Item Status
Parent BF16 checkpoint Public
NVFP4 conversion Passed: ModelOpt NVFP4 W4A4 with BF16 MoVA value experts
Structural verification Passed: index/shards, tensor inventory, required dense tensors, and exhaustive finite-value scan
Full-load smoke Passed on NVIDIA RTX PRO 6000 Blackwell Server Edition
Coherent generation Passed
Reasoning separation Passed
Streaming reasoning/content separation Passed
Native tool-call parsing Passed: offline fixtures and three consecutive live OpenAI API trials
Refusal evaluation Pending
Capability-retention evaluation Pending
GB10 single-request throughput Passed: 21.38 wall-clock completion tok/s median across three runs
Full-window prompt validation Pending; 262,144-token allocation passed
Baked deployment system prompt None planned
Modification/conversion recipe Proprietary and intentionally not distributed
Access Public experimental artifact

DERISKED identifies the Blackfrost research family. It is not a claim of zero refusals, complete safety, harmlessness, lossless quantization, benchmark parity, or production readiness.

Overview

This repository contains a Blackfrost NVFP4 deployment derivative produced from Blackfrost-AI/K2-Horizon-MoVA-36B-A4B-DERISKED-BF16. The parent checkpoint is itself a weight-level behavioral derivative of IFM/K2-Horizon-MoVA-36B-A4B.

The package retains the native K2 Horizon tokenizer, chat template, long-context configuration, custom architecture code, MoE layout, and MoVA layout while representing eligible weights in a ModelOpt-compatible NVFP4 deployment format. K2 Horizon's self_attn.v_experts matrices remain BF16 because the native fused-MoVA runtime consumes dense value-expert matrices; routers, embeddings, normalization tensors, and the LM head also retain their runtime-appropriate higher precision.

Internal behavioral directions, capture data, target maps, calibration material, intermediate checkpoints, evaluation prompts, and reproduction recipes will not be distributed.

Model specifications

Property Value
Parent architecture K2HorizonForCausalLM / k2_horizon
Parameters 36B total / 4B active per token, per the upstream card
Transformer layers 48
Hidden size 2,560
Routed experts 100 total / 8 active per token
Shared experts 1 per MoE layer
MoVA value experts 64 total / 4 active per token
Native maximum context 524,288 tokens
Deployment format Mixed ModelOpt NVFP4 W4A4 / BF16 MoVA value experts, with FP8 KV-cache metadata
ModelOpt source 0.48.0dev tag at 022767c7ab3d7d36211affd85e5c496770cde768
Weight shards 4
Safetensors bytes 33,794,653,520
Tensor inventory 58,515 total: 3,159 BF16; 27,678 FP32; 13,839 FP8 E4M3FN; 13,839 UINT8
Dense MoVA value-expert weights 2,880 BF16 tensors
Conversion smoke runtime vLLM 0.27.1+aeon.sm120.rtx; PyTorch 2.13.0+cu130; Transformers 5.14.1
Spark deployment runtime vLLM 0.1.dev20073+g8e685d198; PyTorch 2.13.0+cu130; Transformers 5.15.1
Validated hardware NVIDIA RTX PRO 6000 Blackwell Server Edition, 96 GB; NVIDIA GB10 Spark
Spark served window 262,144 tokens configured and allocated; full-window prompt validation pending

The native context limit is not a guarantee that the eventual runtime or hardware configuration can allocate or serve the full window.

Lineage

IFM/K2-Horizon-MoVA-36B-A4B@7730b92d1b574e04663b04023d5d6fa83475432f
└── Blackfrost-AI/K2-Horizon-MoVA-36B-A4B-DERISKED-BF16
    └── Blackfrost-AI/K2-Horizon-MoVA-36B-A4B-DERISKED-NVFP4

Prompting and runtime compatibility

The artifact retains the native K2 Horizon chat template, with reasoning_effort="high", temperature=1.0, and top_p=0.95 as the upstream starting settings. No Blackfrost, Frosty, compliance, or deployment-specific system prompt is baked into the package.

The exact artifact passed native K2 Horizon reasoning separation, streaming reasoning/content separation, generic OpenAI multi-turn history, and repeated schema-valid tool-call parsing through an OpenAI-compatible vLLM endpoint. The deployment compatibility template defaults tool calls to K2's native JSON representation; clients may still request K2 XML or typed XML explicitly. Generic NVFP4 support alone does not imply support for K2 Horizon custom code, expert routing, MoVA, parsers, streaming, or long-context allocation.

GB10 OpenAI API deployment kit

The repository includes a self-contained NVIDIA GB10 vLLM deployment kit with:

  • the exact prepacked-MoVA model adapter used for the accepted run;
  • K2 reasoning and tool-call parsers;
  • the OpenAI multi-turn compatibility chat template;
  • an immutable validated ARM64 container reference;
  • a portable launch.sh with no machine-specific paths;
  • a standard-library API verifier for reasoning and tool calls.

The patch removes a decode-time reconstruction of all 64 BF16 MoVA value experts in each of 45 sparse layers. On the accepted GB10 test, median wall-clock completion throughput improved from 6.06 to 21.38 tokens/s, or 3.53x. These rates include prefill, time to first token, detokenization, and local HTTP transport.

Validation record and remaining gates

Completed on the exact artifact:

  1. converter, ModelOpt, CUDA, framework, driver, runtime, and hardware recorded;
  2. exact quantization map, four-shard inventory, byte count, dtype inventory, and exhaustive finite-value scan;
  3. required dense MoVA tensors and native tokenizer/template/custom-code assets preserved;
  4. full native runtime load and coherent generation;
  5. exact single served-model identity and separated reasoning output;
  6. streaming reasoning/content separation with clean termination;
  7. native tool-call parser fixtures and three consecutive live calls with valid JSON arguments;
  8. three-run GB10 throughput benchmark at 21.38 wall-clock completion tok/s median.

Still pending:

  1. refusal and red-line release gates;
  2. capability-retention comparison against the exact parent BF16 checkpoint;
  3. full-window prompt validation and concurrency saturation under representative traffic.

No upstream or BF16-parent benchmark result may be attributed to this quantized derivative without rerunning the exact artifact.

Access

This NVFP4 repository is a public experimental artifact. Public availability does not imply production readiness or completion of the pending behavioral and full-window gates.

Limitations and responsibility

  • This is an experimental quantized research checkpoint.
  • Quantization can alter capability, numerical behavior, refusal behavior, formatting, routing, and long-context stability.
  • Generated content may be inaccurate, insecure, offensive, or otherwise unsuitable.
  • The model is not a security boundary, policy engine, authorization mechanism, or substitute for professional judgment.
  • Operators are responsible for access controls, monitoring, legal compliance, and safeguards appropriate to their environment.

License and attribution

The upstream K2 Horizon repository identifies the model license as Apache-2.0. Review the upstream model card and the BF16 parent repository before use or redistribution.

Blackfrost is independent of and is not affiliated with, sponsored by, or endorsed by IFM. This artifact is provided as-is, without warranties.

Contact

For reproducible artifact issues, use this repository's Discussions. Do not post credentials, private prompts, personal information, or infrastructure details.