Qwen3.5-9B-C3SM-SDM-Agentic-Coder (GGUF)
Official GGUF quantizations for OliviaRossi/Qwen3.5-9B-C3SM-SDM-Agentic-Coder, a state-of-the-art 9B-parameter autonomous coding and multi-turn tool-calling model.
Synthesized using Curvature-Calibrated Spectral Consensus Merging with Sinusoidal Depth Modulation (C3SM-SDM), all multimodal and speculative MTP layers were surgically pruned to yield a pure, clean 32-block hybrid causal LM compatible with downstream llama.cpp runtimes and quantization schemes (Q4_K_M, Q5_K_M, Q8_0).
Model Lineage & Heritage
The model unifies a top-tier distillation anchor with four specialized domain-expert fine-tunes:
| Role | Repository | Domain Specialization |
|---|---|---|
| Anchor Base ($M_0$) | XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B |
Reasoning distillation backbone, native instruction compliance, multi-turn state harness. |
| Donor 1 ($M_1$) | ornith-ai/Ornith-1.5-9B |
Step-by-step logic, conversational reasoning, multi-turn format compliance. |
| Donor 2 ($M_2$) | Jackrong/Qwopus3.5-9B-Coder |
Claude-Opus style agentic coding, deep tool-use schemas, recursive refactoring. |
| Donor 3 ($M_3$) | OrionLLM/OxCoder-9B |
Complex algorithmic synthesis, long-context programming, MCP tool execution. |
| Donor 4 ($M_4$) | empero-ai/Qwen3.8-9B-Distill |
Distilled Qwen3.8 foundation scale, high-capacity general world and syntax intelligence. |
| Chat Template | OliviaRossi/Improved-Chat-Template-for-Qwen-3.x |
Dual-format agentic execution (JSON/XML), multi-tier <thought> reasoning CoT. |
Mathematical Architecture: C3SM-SDM
Standard weight merging algorithms (SLERP, TIES, DARE) apply static global scalars across the entire network, causing representation drift, activation saturation, and degraded token logits. C3SM-SDM merges across four continuous harmonic waveforms:
Normalized Layer Depth Coordinate: z = l / 31 (for layer l in [0..31])
z = 0.0 z = 0.5 z = 1.0
[Shallow Syntactic] ───> [Core Reasoning / Tools] ───> [Output Calibration]
α(0) = 0.20 α(0.5) = 0.85 α(1) = 0.20
p(0) = 0.50 (High DARE) p(0.5) = 0.30 (Dense Delta) p(1) = 0.50 (High DARE)
r(0) = 32 (Low Rank) r(0.5) = 96 (Rich Spectrum) r(1) = 32 (Low Rank)
1. Harmonic Task Amplitude Modulation: $\alpha(z)$
To prevent boundary degradation (token embedding corruption at layer 0 and logit distribution distortion at layer 31), the task arithmetic scale follows a harmonic envelope: $$\alpha(z) = \alpha_{\min} + (\alpha_{\max} - \alpha_{\min}) \cdot \sin^2(\pi z), \quad \alpha \in [0.20, 0.85]$$ The core algorithmic manifold ($z \in [0.35, 0.75]$) receives the maximum fine-tuned coding delta, while syntactic boundaries remain securely anchored to the distilled base.
2. Phase-Shifted Domain Specialization: $\vec{\lambda}_k(z)$
Each donor model is routed according to its structural layer specialty: $$w_k(z) = w_{k, \text{base}} \cdot \left[1.0 + A_k \cdot \sin(\pi z + \phi_k)\right], \quad \lambda_k(z) = \frac{w_k(z)}{\sum_j w_j(z)}$$
- Ornith-1.5 ($\phi = -0.20\pi$): Peaks in early reasoning layers ($z \approx 0.35$) to shape planning and CoT trajectories.
- Qwopus-Coder ($\phi = 0.00\pi$): Centers at mid-depth ($z \approx 0.50$) to govern agentic tool schemas and state transitions.
- OxCoder ($\phi = +0.25\pi$): Centers in deep reasoning blocks ($z \approx 0.65$) for heavy AST transformations and algorithm synthesis.
- Qwen3.8-Distill ($\phi = +0.05\pi$): Provides continuous foundational regularization across all depths.
3. Inverted Harmonic DARE Sparsification: $p(z)$
Parameter noise from fine-tuning collisions is most harmful to syntactic parsing at the boundaries. DARE dropout is inverted: $$p(z) = p_{\max} - (p_{\max} - p_{\min}) \cdot \sin(\pi z), \quad p \in [0.30, 0.50]$$
4. Dynamic Spectral Subspace Denoising (rSVD): $r(z)$
Consensus deltas undergo randomized singular value decomposition to eliminate high-entropy orthogonal interference: $$\Delta_{\text{consensus}} \approx \mathbf{U}{r(z)} \mathbf{\Sigma}{r(z)} \mathbf{V}_{r(z)}^T, \quad r(z) = \text{round}\left(32 + 64 \cdot \sin(\pi z)\right)$$
5. Frobenius Manifold Calibration
To ensure activation stability across GQA and Gated DeltaNet recurrent kernels, the merged weight tensor is projected back to the convex norm of its contributors: $$\mathbf{W}^* = (\mathbf{W}0 + \alpha(z) \Delta{\text{spectral}}) \cdot \frac{(1-\beta)|\mathbf{W}_0|_F + \beta \sum \lambda_k |\mathbf{W}_k|F}{|\mathbf{W}{\text{merged}}|_F}$$
Quantization Matrix & Hardware Guidance
All quants are packaged as single-file GGUF binaries converted directly from the 32-block causal architecture:
| File Name | Precision | File Size | Recommended VRAM / RAM | Primary Use Case |
|---|---|---|---|---|
Qwen3.5-9B-C3SM-SDM-Agentic-Coder-Q5_K_M.gguf |
5-bit Medium | ~6.4 GiB | 12 GB VRAM (RTX 3060 12GB, RTX 4070) | Near-Lossless Precision: Superior syntax retention, complex multi-file refactoring, and strict schema adherence. |
Chat Template & Agent Prompting
The GGUF metadata includes the Olivia Rossi Improved Chat Template for Qwen 3.x. It natively parses structured tool calls, system-level declarations, and multi-tier Chain-of-Thought reasoning.
Tool Call Format (OpenAI / JSON Compatible)
<|im_start|>system
You are an expert autonomous coding agent.
<tools>
[
{
"name": "execute_bash",
"description": "Execute a shell command in the workspace environment.",
"parameters": {
"type": "object",
"properties": {
"command": {"type": "string", "description": "The command to run."}
},
"required": ["command"]
}
}
]
</tools><|im_end|>
<|im_start|>user
Inspect the current directory and write a Python test runner script.<|im_end|>
<|im_start|>assistant
<thought>
The user wants to inspect the repository structure before implementing the test runner.
First step: list the directory contents using the bash tool.
</thought>
<tool_call>
{"name": "execute_bash", "arguments": {"command": "ls -la"}}
</tool_call><|im_end|>
<|im_start|>tool
<tool_response>
total 16
drwxr-xr-x 2 user user 4096 Sep 27 12:00 src
-rw-r--r-- 1 user user 420 Sep 27 12:00 pyproject.toml
</tool_response><|im_end|>
<|im_start|>assistant
<thought>
Directory contains `src/` and `pyproject.toml`. Now generate the pytest execution script.
</thought>
I have inspected the project directory. Let's create `run_tests.py`...<|im_end|>
Quickstart & Deployment
1. In Ollama
Create a Modelfile:
FROM ./Qwen3.5-9B-C3SM-SDM-Agentic-Coder-Q5_K_M.gguf
PARAMETER temperature 0.2
PARAMETER top_p 0.95
PARAMETER num_ctx 32768
PARAMETER stop "<|im_end|>"
PARAMETER stop "<|endoftext|>"
Build and launch:
ollama create qwen3.5-agentic -f Modelfile
ollama run qwen3.5-agentic
2. In llama-server (OpenAI-Compatible Local Endpoint)
To run as an autonomous agent backend for Aider, Cline, or Roo Code:
./llama-server \
-m Qwen3.5-9B-C3SM-SDM-Agentic-Coder-Q5_K_M.gguf \
-c 32768 \
-ngl 99 \
--host 0.0.0.0 \
--port 8080
Configure your IDE client:
- API Base:
http://localhost:8080/v1 - Model ID:
Qwen3.5-9B-C3SM-SDM-Agentic-Coder-Q5_K_M - API Key:
sk-no-key-required
3. In llama-cli
./llama-cli \
-m Qwen3.5-9B-C3SM-SDM-Agentic-Coder-Q4_K_M.gguf \
-p "<|im_start|>system\nYou are a helpful coding assistant.<|im_end|>\n<|im_start|>user\nWrite a lock-free multi-producer single-consumer ring buffer in Rust.<|im_end|>\n<|im_start|>assistant\n<thought>\n" \
-n 2048 \
-c 8192 \
-ngl 99
Architectural Verification
- Decoder Layers: 32 regular transformer blocks (
blk.0throughblk.31). - Linear Attention: Gated DeltaNet hybrid recurrence preserved with zero dropout perturbation.
- Multimodal Projectors: Stripped (
visual.*,vision_tower.*,mmprojremoved). - MTP Heads: Speculative decoding layers eliminated for clean, zero-error GGUF load validation.
- Context Capacity: Up to 262,144 tokens supported via native RoPE scaling.
Citation & Acknowledgments
If you utilize this model or the C3SM-SDM merging formulation in your research or applications, please cite:
@misc{rossi2026c3smsdm,
author = {Olivia Rossi},
title = {Qwen3.5-9B-C3SM-SDM-Agentic-Coder: Curvature-Calibrated Spectral Consensus Merging with Sinusoidal Depth Modulation},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/OliviaRossi/Qwen3.5-9B-C3SM-SDM-Agentic-Coder-GGUF}}
}
Special gratitude to the authors of Xiaomi MiMo, Ornith AI, Jackrong (Qwopus), OrionLLM (OxCoder), and Empero AI for their open-weights contributions to the open-source coding ecosystem.