npario/Ornith-1.5-35B-A3B-6bit-XL-mlx

🤗 Hugging Face 来源image-text-to-textmit35.1B 参数激活 3B70 GBsafetensors✓ 7 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo npario/Ornith-1.5-35B-A3B-6bit-XL-mlx ./model-folder
需要做种者 →

leonsarmiento/Ornith-1.5-35B-A3B-6bit-XL-mlx

Ornith-1.5-35B-A3B by Ornith AI — a 35B-A3B sparse MoE built through end-to-end self-improvement (jointly optimized task generation, scaffold construction, and solution rollouts via RL), quantized for Apple Silicon using the BaseQuant_XL 6/8-bit recipe.

Only ~3B parameters active per token — the decode speed of a small model with the capacity of a 35B one. Ornith-1.5 is an agentic-coding specialist: 79 SWE-bench Verified, 59.6 SWE-bench Pro, 67.8/68.5 Terminal-Bench 2.1 — beating its Ornith-1.0 parent (+3.4 SWE-bench) and Qwen3.6-35B-A3B (+5.6) by wide margins.

This is a full multimodal build — the vision tower is preserved.

About XL Quantization

BaseQuant_XL is a fully data-agnostic, static quantization. No calibration dataset, no sensitivity analysis, no importance matrix. Precision is allocated purely by architectural role — routing-critical layers get higher precision, bulk expert parameters get lower precision. The result is a transparent, faithful capture of the source model.

Data-dependent calibration quantizations (iMatrix, AWQ, GPTQ, oQ, oQ4e, etc.) use a calibration set to guide bit allocation. This can produce a skewed representation of the model: domains well-represented in the calibration data (English, popular topics, public or leaked benchmarks) are preserved better, while underrepresented domains (non-English languages, niche use cases, your own data) are preserved worse. XL avoids this trade-off entirely — it generalizes honestly because it is never fit to any particular data distribution.

Quickstart

pip install -U mlx-vlm
python -m mlx_vlm.generate --model leonsarmiento/Ornith-1.5-35B-A3B-6bit-XL-mlx --max-tokens 256 --temperature 0.6 --top-p 0.95 --prompt "Implement an LRU cache in Python with O(1) get/put."

Works with LM Studio — vision mmproj included. Thinking mode is on by default (emits <think>...</think>).

Quantization Strategy

BaseQuant_XL recipe — precision is allocated by layer importance, not applied uniformly:

Layers Bits Rationale
mlp.gate (router), shared_expert_gate, lm_head, shared_expert bf16 Routing decisions and output projection — any quantization noise here causes expert misrouting or output degradation
embed_tokens, self_attn, linear_attn 8-bit Every-token layers — near-lossless, attention quality preserved
vision_tower, switch_mlp (routed experts) 6-bit Bulk parameters — 256 experts with only 8 active per token; redundancy absorbs quantization noise. 6-bit is the sweet spot for routed experts (higher bits can cause overthinking)
  • Bits per weight: ~6.8 · Total size: ~28 GB · Group size: 64

Notes specific to this build:

  • Source ships an MTP (multi-token prediction) layer — dropped in this build (engines that execute bundled MTP are the exception, not the rule; standard XL builds omit it).
  • Unlike Ornith-1.0 (which stores experts individually), Ornith-1.5's source stores merged per-layer expert tensors — converts via stock mlx_vlm sanitize.

Recommended Inference Parameters

Per the source model card — new settings for the Ornith-1.5 generation:

Parameter Value
temperature 0.6 (general tasks) · 1.0 (reproducing reported benchmarks)
top_p 0.95
top_k 20
presence_penalty 1.1
reasoning_parser qwen3 (changed from deepseek_r1 in Ornith-1.0)
tool_call_parser qwen3_xml

Thinking is on by default (<think>...</think> before the answer); with a reasoning parser enabled the chain-of-thought is returned in a separate reasoning_content field.

Model Overview

Property Value
Architecture Qwen3.5-family MoE (35B-A3B) + native vision encoder
Parameters 35.9B total / ~3B active per token
Experts 256 (8 routed + 1 shared)
Attention Hybrid — 30 linear_attn + 10 full attention (40 layers)
Modalities text, image, video → text
Context window 262,144 tokens native
Thinking <think>...</think>
Tool calling XML-style (<tool_call><function=...><parameter=...>)
License MIT

Source Model Benchmarks (from Ornith AI)

Benchmark Ornith-1.5-35B-A3B Ornith-1.0-35B-A3B Qwen3.6-35B-A3B
SWE-bench Verified 79.0 75.6 73.4
SWE-bench Pro 59.6 50.4 49.5
Terminal-Bench 2.1 (Terminus-2) 67.8 64.2 52.5
Terminal-Bench 2.1 (Claude Code) 68.5 62.8 49.2

Source

ornith-ai/Ornith-1.5-35B-A3B — Ornith 1.5 blog post