Gemory-26B-A4B - GGUF
This repository contains GGUF quantizations of UltimateIntent/Gemory-26B-A4B-GGUF.
Gemory-26B-A4B is a 26-billion parameter Mixture of Experts (MoE) model built on the Gemma architecture foundation, fine-tuned for high-fidelity creative writing, unrestricted roleplay, and complex multi-turn conversational depth. All quantizations were generated directly from the uncompressed 50.5 GB BF16 base model and sanity-tested for generation integrity.
Quantization Breakdown
| Filename | Quant Method | File Size | Recommended RAM/VRAM | Description / Quality |
|---|---|---|---|---|
Gemory-26B-A4B-Q6_K.gguf |
Q6_K | 22.6 GB | ~26 GB+ | Near-lossless quality; closest output to base BF16 weights. |
Gemory-26B-A4B-Q5_K_M.gguf |
Q5_K_M | 19.1 GB | ~23 GB+ | High quality; recommended sweet spot for 24 GB VRAM GPUs. |
Gemory-26B-A4B-Q5_K_S.gguf |
Q5_K_S | 18.0 GB | ~22 GB+ | High quality with slightly faster inference and lower memory. |
Gemory-26B-A4B-Q4_K_M.gguf |
Q4_K_M | 16.8 GB | ~20 GB+ | Default Recommended: Optimal balance of speed, perplexity, and memory. |
Gemory-26B-A4B-Q4_K_S.gguf |
Q4_K_S | 15.5 GB | ~19 GB+ | Standard 4-bit quantization for tighter memory envelopes. |
Gemory-26B-A4B-Q3_K_M.gguf |
Q3_K_M | 13.3 GB | ~16 GB+ | Low-memory 3-bit variant; preserves critical attention weights. |
Gemory-26B-A4B-Q3_K_S.gguf |
Q3_K_S | 12.2 GB | ~15 GB+ | Maximum compression; fits easily on 16 GB systems. |
Usage Instructions
1. llama.cpp
Command-Line Inference (llama-cli):
./llama-cli \
-hf Abiray/Gemory-26B-A4B-GGUF:Q4_K_M \
-p "Write an opening scene set in a rain-soaked neon alleyway." \
-n 512 \
--temp 0.7