WhiskyAKM/G9v3-3B-NVFP4-GGUF

🤗 Hugging Face sourcetext-generationapache-2.03B activated1.8 GBGGUF✓ 1 checksumupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo WhiskyAKM/G9v3-3B-NVFP4-GGUF ./model-folder
Needs a seeder →

G9v3-3B NVFP4 GGUF

NVFP4 (4-bit floating-point) quantized version of ai9stars/G9v3-3B, a compact 3B-parameter Llama-architecture language model supporting English and Chinese, with a 128K context window, tool-calling, and a built-in thinking/reasoning mode.

Model Overview

G9v3-3B is a lightweight text-generation model built on the Llama architecture. Despite its small size (~3B parameters), it supports a 131,072-token context window (128K) thanks to a high RoPE θ of 5,000,000. The model uses Grouped-Query Attention (GQA) with 2 key-value heads across 16 attention heads for efficient inference, and includes special tokens for thinking/reasoning (<|thought_begin|> / <|thought_end|>, /think, /no_think) and tool calling (<function>, <|tool_call|>, etc.).

The model uses a ChatML-style conversation format with <|im_start|> / <|im_end|> delimiters.

Model Architecture

Property Value
Architecture Llama
Parameters ~3B
Hidden Size 2048
Intermediate Size 6144
Layers 52
Attention Heads 16
KV Heads 2 (GQA)
Head Dimension 128
Context Length 131,072
Vocabulary Size 130,560
RoPE Theta 5,000,000
Original Precision bfloat16
Supported Languages en, zh

Quantization

This model is quantized with NVFP4 (NVIDIA 4-bit floating point). NVFP4 is NVIDIA's 4-bit floating-point format that preserves a wider dynamic range than integer 4-bit formats (such as Q4_K), offering better accuracy while remaining highly memory-efficient.

File Quantization Size Precision
g9v3-3b-nvfp4.gguf NVFP4 1.7 GB 4-bit float

Usage

llama.cpp CLI

./llama-cli \
  -m g9v3-3b-nvfp4.gguf \
  -p "Explain quantum computing in simple terms." \
  --temp 0.9 --top-p 0.95

llama-server (OpenAI-compatible API)

./llama-server \
  -m g9v3-3b-nvfp4.gguf \
  --host 0.0.0.0 --port 8080

Thinking Mode

The model supports a thinking/reasoning mode controlled via special tokens. Use /think to enable extended reasoning or /no_think to disable it. When thinking is enabled, the model outputs its reasoning between <|thought_begin|> and <|thought_end|> tokens before providing the final answer.

Tool Calling

The model supports function/tool calling via XML-style <function> tags. Tool definitions are injected into the system prompt, and the model responds with <function name="..."> blocks containing <param> elements.

Generation Parameters

Recommended parameters from the original model's generation_config.json:

Parameter Value
Temperature 0.9
Top-P 0.95
BOS Token ID 0
EOS Token IDs 1, 130073
Pad Token ID 1

Acknowledgements

License

Apache License 2.0