Holo4-35B-A3B GGUF
GGUF quantizations of Hcompany/Holo4-35B-A3B, a 35B mixture-of-experts (MoE, 3B active) vision-language model (VLM) for Computer Use, tool-driven work, and agentic workflows.
For image input, use the included F16 vision projector (mmproj-Holo4-35B-A3B-f16.gguf).
Upstream benchmarks
Results reported by H Company from evaluations of the original Holo4-35B-A3B model across computer tasks, long workflows, and tool servers.
GGUF files
| Quantization | File | Size | Notes |
|---|---|---|---|
| Q4_0 | Holo4-35B-A3B-Q4_0.gguf | 18.4 GB | Standard Q4 0-quant |
| Q4_K_M | Holo4-35B-A3B-Q4_K_M.gguf | 19.7 GB | Recommended balanced default |
| Q5_K_M | Holo4-35B-A3B-Q5_K_M.gguf | 23.0 GB | Higher-quality Q5 option |
| Q6_K | Holo4-35B-A3B-Q6_K.gguf | 26.6 GB | High-quality option |
| Q8_0 | Holo4-35B-A3B-Q8_0.gguf | 34.4 GB | Near-lossless reference quantization |
| Vision projector | mmproj-Holo4-35B-A3B-f16.gguf | 858 MB | Required for image input |
Usage
Use a current llama.cpp build with the included chat template.
llama-cli \
-m Holo4-35B-A3B-Q4_K_M.gguf \
-c 4096 -n 512 --temp 0.7 --top-p 0.8 \
--jinja --chat-template-file chat_template.jinja \
-p "<|im_start|>user\nWhat is computer use in AI?<|im_end|>\n<|im_start|>assistant\n"
For multimodal / image input:
llama-mtmd-cli \
-m Holo4-35B-A3B-Q4_K_M.gguf \
-mm mmproj-Holo4-35B-A3B-f16.gguf \
--image screenshot.png \
-p "Describe the interface in this screenshot."
Source
- Model: Hcompany/Holo4-35B-A3B
- Chat template revision:
5084587 - License: apache-2.0
- Checksums: SHA256SUMS