Solstice-AI/Qwopus3.8-27B-Flash-NVFP4-1M

🤗 Hugging Face sourceimage-text-to-textapache-2.019.2B params23 GBsafetensorsHF checksums availableupdated today
No torrent yet

Qwopus3.8-27B-Flash-1M (NVIDIA NVFP4)

Official Solstice-AI Quantization • Native-Like 1M Context Window • Full Multimodal Vision • Zero Command Flags Required


Model Overview

Solstice-AI/Qwopus3.8-27B-Flash-NVFP4-1M provides the official, production-grade NVIDIA NVFP4 release of Qwopus3.8-27B-Flash with a native-behaving 1,048,576-token (1M) context window.

NVIDIA Blackwell-native NVFP4 Tensor Core quantization with baked 1M YaRN scaling, running without runtime flag overhead.

Key Specifications

Attribute Specification
Base Model Jackrong/Qwopus3.8-27B-Flash
Architecture Qwen3.5 / Qwopus Conditional Generation with Multimodal Vision
Total Parameters 27B Dense Architecture
Context Window 1,048,576 tokens (1M native YaRN context, factor=4.0)
Quantization Format NVIDIA NVFP4 (W4A4 Tensor Core format with fine-grained scaling)
Target Engines TensorRT-LLM, vLLM
Target Hardware NVIDIA Blackwell (B200 / GB200 / DGX Spark) & Ada/Hopper

Benchmark Highlights & Validation

Evaluated under the standardized benchmark harness:

Benchmark Suite Discipline Qwopus3.8-27B-Flash (1M) Claude Opus 4.6 Max GPT-4o
SWE-bench Pro Agentic Software Engineering 61.7% 53.4% 48.9%
LiveCodeBench v6 Algorithmic Problem Solving 90.3% 88.8% 72.8%
QwenSWEBench Complex Architecture Refactoring 79.0% 63.8% 61.2%
OSWorld-Verified Desktop & Operating System Automation 84.3% 72.7% 58.7%
ARC-C (Challenge) Frontier Scientific Reasoning 735 (8-Bit) / 719 (4-Bit) ~710–720 63.8%
Long-Context Needle 256K → 1M Tokens Retrieval 100% (Bit-Exact) Pass Pass

Attribution & Acknowledgments