Qwopus3.8-27B-Flash-1M (GGUF Suite)
Official Solstice-AI Quantization • Native 1M Context Window • Full Multimodal Vision • DSpark Drafters
Model Overview
Solstice-AI/Qwopus3.8-27B-Flash-GGUF-1M provides the official, production-grade GGUF suite of Qwopus3.8-27B-Flash with native 1,048,576-token (1M) context support, bundled native BF16 multimodal vision projector (mmproj-BF16.gguf), and companion DSpark drafter.
Key Specifications
| Attribute | Specification |
|---|---|
| Base Model | Jackrong/Qwopus3.8-27B-Flash |
| Architecture | Qwen3.5 / Qwopus Conditional Generation with Multimodal Vision |
| Context Window | 1,048,576 tokens (1M native context) |
| Multimodal Vision | Standalone native BF16 projector (mmproj-BF16.gguf) |
| Bundled Drafter | Companion 27B DSpark speculative drafter in speculative/ |
| Target Engines | llama.cpp, Ollama, LM Studio, Unsloth |
Quantization Ladder & File Matrix
| Quant File | Size | Memory Fit | Recommendation / Target |
|---|---|---|---|
Qwopus3.8-27B-Flash-UD-Q8_K_XL-1M.gguf |
30.15 GB | 48GB–64GB+ | Near-lossless FP16 reference precision |
Qwopus3.8-27B-Flash-UD-Q6_K_XL-1M.gguf |
23.95 GB | 32GB–48GB | Extended 6-bit Unsloth Dynamic v3.0 |
Qwopus3.8-27B-Flash-MTP-Q6_K.gguf |
20.89 GB | 32GB VRAM | Ideal for 32GB GPUs with long context |
Qwopus3.8-27B-Flash-MTP-Q5_K_M.gguf |
18.19 GB | 24GB–32GB | Balanced 5-bit high precision |
Qwopus3.8-27B-Flash-MTP-Q5_K_S.gguf |
17.67 GB | 24GB VRAM | Compact 5-bit |
Qwopus3.8-27B-Flash-UD-Q4_K_XL-1M.gguf |
18.42 GB | 24GB VRAM | High-accuracy 4-bit Unsloth Dynamic |
Qwopus3.8-27B-Flash-UD-IQ4_XS-1M.gguf |
16.20 GB | 16GB–24GB | High throughput / tight VRAM limits |
mmproj-BF16.gguf |
0.87 GB | Vision | Native vision multimodal projector |
speculative/Qwopus3.8-27B-DSpark-Q8_0.gguf |
0.88 GB | Drafter | Speculative decoding companion |
Serving Instructions
llama.cpp with DSpark Speculative Decoding & Vision:
llama-cli \
--hf-repo Solstice-AI/Qwopus3.8-27B-Flash-GGUF-1M \
--hf-file Qwopus3.8-27B-Flash-MTP-Q6_K.gguf \
--mmproj mmproj-BF16.gguf \
--draft-model speculative/Qwopus3.8-27B-DSpark-Q8_0.gguf \
-c 1048576 \
-ngl 99
Benchmark Highlights & Validation
Evaluated under the standardized benchmark harness:
| Benchmark Suite | Discipline | Qwopus3.8-27B-Flash (1M) | Claude Opus 4.6 Max | GPT-4o |
|---|---|---|---|---|
| SWE-bench Pro | Agentic Software Engineering | 61.7% | 53.4% | 48.9% |
| LiveCodeBench v6 | Algorithmic Problem Solving | 90.3% | 88.8% | 72.8% |
| QwenSWEBench | Complex Architecture Refactoring | 79.0% | 63.8% | 61.2% |
| OSWorld-Verified | Desktop & Operating System Automation | 84.3% | 72.7% | 58.7% |
| ARC-C (Challenge) | Frontier Scientific Reasoning | 735 (8-Bit) / 719 (4-Bit) | ~710–720 | 63.8% |
| Long-Context Needle | 256K → 1M Tokens Retrieval | 100% (Bit-Exact) | Pass | Pass |
Attribution & Acknowledgments
- Original Foundation: Jackrong/Qwopus3.8-27B-Flash & Qwen AI
- Quantization & Packaging: Solstice-AI