Abiray/Gemory-26B-A4B-GGUF

🤗 Hugging Face 来源text-generationapache-2.0激活 4B118 GBGGUF✓ 7 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo Abiray/Gemory-26B-A4B-GGUF ./model-folder
需要做种者 →

Gemory-26B-A4B - GGUF

This repository contains GGUF quantizations of UltimateIntent/Gemory-26B-A4B-GGUF.

Gemory-26B-A4B is a 26-billion parameter Mixture of Experts (MoE) model built on the Gemma architecture foundation, fine-tuned for high-fidelity creative writing, unrestricted roleplay, and complex multi-turn conversational depth. All quantizations were generated directly from the uncompressed 50.5 GB BF16 base model and sanity-tested for generation integrity.


Quantization Breakdown

Filename Quant Method File Size Recommended RAM/VRAM Description / Quality
Gemory-26B-A4B-Q6_K.gguf Q6_K 22.6 GB ~26 GB+ Near-lossless quality; closest output to base BF16 weights.
Gemory-26B-A4B-Q5_K_M.gguf Q5_K_M 19.1 GB ~23 GB+ High quality; recommended sweet spot for 24 GB VRAM GPUs.
Gemory-26B-A4B-Q5_K_S.gguf Q5_K_S 18.0 GB ~22 GB+ High quality with slightly faster inference and lower memory.
Gemory-26B-A4B-Q4_K_M.gguf Q4_K_M 16.8 GB ~20 GB+ Default Recommended: Optimal balance of speed, perplexity, and memory.
Gemory-26B-A4B-Q4_K_S.gguf Q4_K_S 15.5 GB ~19 GB+ Standard 4-bit quantization for tighter memory envelopes.
Gemory-26B-A4B-Q3_K_M.gguf Q3_K_M 13.3 GB ~16 GB+ Low-memory 3-bit variant; preserves critical attention weights.
Gemory-26B-A4B-Q3_K_S.gguf Q3_K_S 12.2 GB ~15 GB+ Maximum compression; fits easily on 16 GB systems.

Usage Instructions

1. llama.cpp

Command-Line Inference (llama-cli):

./llama-cli \
  -hf Abiray/Gemory-26B-A4B-GGUF:Q4_K_M \
  -p "Write an opening scene set in a rain-soaked neon alleyway." \
  -n 512 \
  --temp 0.7