WhiskyAKM/Gemma-4-E2B-it-qat-NVFP4-GGUF

🤗 Hugging Face sourceany-to-anyapache-2.07.0 GBGGUF✓ 3 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo WhiskyAKM/Gemma-4-E2B-it-qat-NVFP4-GGUF ./model-folder
Needs a seeder →

Gemma 4 E2B IT QAT - NVFP4 GGUF

NVFP4 GGUF quantization based on WhiskyAKM/Gemma-4-E2B-it-qat-GGUF, created from the Quantization-Aware Training (QAT) checkpoint of Gemma 4 E2B IT.

Model Overview

Gemma 4 E2B IT is a multimodal model built by Google DeepMind that handles text, image, and audio inputs and generates text output. It is designed for efficient on-device deployment on laptops and mobile devices.

The "E" in E2B stands for "effective" parameters - the model uses Per-Layer Embeddings (PLE) to maximize parameter efficiency, giving each decoder layer its own small embedding for every token. This keeps the effective parameter count much smaller than the total.

This repository contains an NVFP4 GGUF conversion of the QAT-optimized checkpoint, making it usable with llama.cpp and other GGUF-compatible inference engines.

Model Architecture

Property Value
Architecture Gemma4ForConditionalGeneration
Effective Parameters 2.3B (5.1B with embeddings)
Layers 35
Sliding Window 512 tokens
Context Length 128K tokens
Vocabulary Size 262K
Supported Modalities Text, Image, Audio
Vision Encoder Parameters ~150M
Audio Encoder Parameters ~300M
Attention Hybrid (sliding window + global, every 5th layer)
RoPE Proportional RoPE (p-RoPE) on global layers

GGUF Files

File Format Size Description
gemma-4-E2B-it-qat-nvfp4.gguf NVFP4 3.2G NVFP4 quantized model - embeddings are in Q6_K
gemma-4-E2B-it-qat-fast-nvfp4.gguf NVFP4 2.6G NVFP4 quantized model - embeddings are in NVFP4
mmproj.gguf - 942M Multimodal projector for image and audio inputs

A chat_template.jinja file is also provided for use with chat-based inference.

Note on NVFP4: NVFP4 (NVIDIA 4-bit Floating Point) is a low-precision floating-point format designed for NVIDIA Blackwell and later GPUs. The QAT optimization preserves quality close to bfloat16 while dramatically reducing memory requirements.

Usage

llama.cpp (CLI)

# Run text inference
./llama-cli \
  -m gemma-4-E2B-it-qat-nvfp4.gguf \
  -p "Explain quantum computing in simple terms." \
  --temp 1.0 --top-k 64 --top-p 0.95

llama-server (OpenAI-compatible API)

./llama-server \
  -m gemma-4-E2B-it-qat-nvfp4.gguf \
  --host 0.0.0.0 --port 8080

Multimodal (Image / Audio)

For image and audio inputs, use a GGUF-compatible engine with multimodal support. Refer to your inference engine's documentation for passing image/audio data alongside text prompts.

Modality order tip: For best results, place image content before text and audio content after text in your prompt.

Generation Parameters

Recommended parameters from the model's generation_config.json:

Parameter Value
Temperature 1.0
Top-K 64
Top-P 0.95
BOS Token ID 2
EOS Token IDs 1, 106, 50
Pad Token ID 0

Thinking Mode

Gemma 4 supports configurable thinking (reasoning) mode:

  • Enable: Include the <|think|> token at the start of the system prompt.
  • Output format: When thinking is enabled, the model outputs internal reasoning followed by the final answer:
    <|channel>thought
    [Internal reasoning]
    <channel|>
    [Final answer]
    
  • Disable: Omit the <|think|> token. For the E2B variant, thinking is fully off when disabled (no empty thought block is generated).

Many libraries like Transformers and llama.cpp handle the chat template complexities automatically.

Key Features

  • Multimodal: Text, image, and audio understanding
  • Long Context: 128K token context window
  • Function Calling: Native support for structured tool use (agentic workflows)
  • Multilingual: Support for 140+ languages
  • On-Device Optimized: Designed for efficient local execution on laptops and mobile devices
  • Native System Prompt: Supports the system role for structured conversations

Acknowledgements

Citation

@misc{gemmateam2026gemma4,
      title={Gemma 4 Technical Report},
      author={Gemma Team},
      year={2026},
      eprint={2607.02770},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2607.02770},
}

License

Apache License 2.0