abenzerps/Holo4-35B-A3B-MLX

认证创作者 abenzerps 已认证
🤗 Hugging Face 来源image-text-to-textapache-2.0激活 3B87 GBsafetensors✓ 23 个校验和今天更新
需要做种者 →

Holo4-35B-A3B MLX

MLX quantizations of Hcompany/Holo4-35B-A3B, a 35B mixture-of-experts (MoE, 3B active) vision-language model (VLM) for Computer Use, tool-driven work, and agentic workflows.

Each quantization includes the full multimodal vision tower (333 visual weights + vision_config), enabling native image and screenshot understanding in MLX-VLM, LM Studio, and Apple Silicon workflows.

Upstream benchmarks

Results reported by H Company from evaluations of the original Holo4-35B-A3B model across computer tasks, long workflows, and tool servers.

MLX Files

Quantization File Size
4-bit Holo4-35B-A3B-MLX-4bit 19.00 GB
6-bit Holo4-35B-A3B-MLX-6bit 27.07 GB
8-bit Holo4-35B-A3B-MLX-8bit 35.13 GB

Multimodal Architecture

Component Architecture Precision
Language Backbone Qwen3.5 MoE (35B total, 3B active) Quantized (4/6/8-bit affine, group_size=64)
Vision Tower Vision Transformer (333 weights) Full precision (BF16)
Projector Multimodal cross-attention / MLP Full precision (BF16)

Vision tower weights and multimodal projectors are preserved in full precision (BF16) to ensure optimal visual comprehension, OCR, and GUI element grounding without degradation.

Usage with MLX-VLM

Installation

pip install -U mlx-vlm

Python API

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

# Load 4-bit (or subfolder="Holo4-35B-A3B-MLX-6bit", subfolder="Holo4-35B-A3B-MLX-8bit")
model_path = "abenzerps/Holo4-35B-A3B-MLX"
subfolder = "Holo4-35B-A3B-MLX-4bit"

model, processor = load(model_path, subfolder=subfolder)
config = load_config(model_path, subfolder=subfolder)

prompt = "Describe the user interface elements shown in this screenshot."
image = ["screenshot.png"]

formatted_prompt = apply_chat_template(
    processor, config, prompt, num_images=len(image)
)

output = generate(model, processor, formatted_prompt, image, verbose=True)
print(output)

Command Line Interface

# 4-bit
python -m mlx_vlm.generate \
    --model abenzerps/Holo4-35B-A3B-MLX --subfolder Holo4-35B-A3B-MLX-4bit \
    --image screenshot.png \
    --prompt "What action should be taken next to achieve the user goal?"

# 6-bit
python -m mlx_vlm.generate \
    --model abenzerps/Holo4-35B-A3B-MLX --subfolder Holo4-35B-A3B-MLX-6bit \
    --image screenshot.png \
    --prompt "What action should be taken next to achieve the user goal?"

# 8-bit
python -m mlx_vlm.generate \
    --model abenzerps/Holo4-35B-A3B-MLX --subfolder Holo4-35B-A3B-MLX-8bit \
    --image screenshot.png \
    --prompt "What action should be taken next to achieve the user goal?"

Source