WhiskyAKM/Gemma-4-12B-it-qat-GGUF

🤗 Hugging Face sourceany-to-anyapache-2.012B activated85 GBGGUF✓ 9 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo WhiskyAKM/Gemma-4-12B-it-qat-GGUF ./model-folder
Needs a seeder →

Gemma 4 12B IT QAT — GGUF

GGUF quantizations of google/gemma-4-12B-it-qat-q4-0, created from the Quantization-Aware Training (QAT) checkpoint of Gemma 4 12B IT.

Model Overview

Gemma 4 12B IT is a multimodal model built by Google DeepMind that handles text, image, and audio inputs and generates text output. It is designed for efficient on-device and server deployment.

This repository contains GGUF conversions of the QAT-optimized checkpoint, making it usable with llama.cpp and other GGUF-compatible inference engines.

Model Architecture

Property Value
Architecture Gemma4ForConditionalGeneration
Parameters 12B
Layers 48
Embedding Dimension 3840
Feed Forward Length 15360
Attention Heads 16
Sliding Window 1024 tokens
Context Length 256K tokens (262144)
Vocabulary Size 262K (262144)
Supported Modalities Text, Image, Audio
Attention Hybrid (sliding window + global, every 6th layer)
RoPE Proportional RoPE (p-RoPE) on global layers
Logit Softcapping 30.0

GGUF Files

File Format Size Description
gemma-4-12b-it-qat-Q4_0.gguf Q4_0 6.5G QAT Q4_0 — native QAT quantization
gemma-4-12b-it-qat-Q4_K_M.gguf Q4_K_M 6.9G K-quant, medium
gemma-4-12b-it-qat-Q4_K_S.gguf Q4_K_S 6.6G K-quant, small
gemma-4-12b-it-qat-Q5_K_M.gguf Q5_K_M 8.0G K-quant, medium
gemma-4-12b-it-qat-Q5_K_S.gguf Q5_K_S 7.8G K-quant, small
gemma-4-12b-it-qat-Q6_K.gguf Q6_K 9.2G K-quant, higher precision
gemma-4-12b-it-qat-Q8_0.gguf Q8_0 12G 8-bit, highest GGUF precision
gemma-4-12b-it-qat-bf16.gguf bf16 23G Full bfloat16 (unquantized)
mmproj.gguf — 168M Multimodal projector (vision + audio)

A chat_template.jinja file is also provided for use with chat-based inference.

Note on QAT: The Q4_0 file is the native QAT quantization. The K-quant and Q8_0 variants are additional GGUF quantizations produced from the QAT checkpoint. The QAT optimization preserves quality close to bfloat16 while dramatically reducing memory requirements.

Note on mmproj: The mmproj.gguf file contains the vision and audio projectors needed for multimodal (image/audio) inference. It is shared across all quantization variants.

Usage

llama.cpp (CLI)

# Run text-only inference
./llama-cli \
  -m gemma-4-12b-it-qat-Q4_0.gguf \
  -p "Explain quantum computing in simple terms." \
  --temp 1.0 --top-k 64 --top-p 0.95

llama-server (OpenAI-compatible API)

# Text-only
./llama-server \
  -m gemma-4-12b-it-qat-Q4_0.gguf \
  --host 0.0.0.0 --port 8080

# Multimodal (image + audio)
./llama-server \
  -m gemma-4-12b-it-qat-Q4_0.gguf \
  --mmproj mmproj.gguf \
  --host 0.0.0.0 --port 8080

Multimodal (Image / Audio)

For image and audio inputs, use llama-server or llama-cli with the --mmproj flag pointing to mmproj.gguf. Refer to your inference engine's documentation for passing image/audio data alongside text prompts.

Modality order tip: For best results, place image content before text and audio content after text in your prompt.

Generation Parameters

Recommended parameters from the model's generation_config.json:

Parameter Value
Temperature 1.0
Top-K 64
Top-P 0.95
BOS Token ID 2
EOS Token ID 1
Pad Token ID 0
Mask Token ID 4

Thinking Mode

Gemma 4 supports configurable thinking (reasoning) mode:

  • Enable: Include the <|think|> token at the start of the system prompt.
  • Output format: When thinking is enabled, the model outputs internal reasoning followed by the final answer:
    <|channel>thought
    [Internal reasoning]
    <channel|>
    [Final answer]
    
  • Disable: Omit the <|think|> token.

Many libraries like Transformers and llama.cpp handle the chat template complexities automatically.

Key Features

  • Multimodal: Text, image, and audio understanding
  • Long Context: 256K token context window
  • Function Calling: Native support for structured tool use (agentic workflows)
  • Multilingual: Support for 140+ languages
  • Native System Prompt: Supports the system role for structured conversations

Acknowledgements

Citation

@misc{gemmateam2026gemma4,
      title={Gemma 4 Technical Report},
      author={Gemma Team},
      year={2026},
      eprint={2607.02770},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2607.02770},
}

License

Apache License 2.0