NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF
GGUF quantizations of nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16, converted with llama.cpp for fast local inference on CPU/GPU.
Files
| Quant | Use case |
|---|---|
| F16 | Full precision, reference quality |
| Q8_0 | Near-lossless, largest quant size |
| Q6_K | Very high quality, minimal loss |
| Q5_K_M / Q5_K_S | High quality, good balance |
| Q4_K_M / Q4_K_S | Recommended default — best speed/quality tradeoff |
| Q4_0 | Legacy 4-bit, faster on some hardware |
| Q3_K_L / Q3_K_M / Q3_K_S | Lower RAM, noticeable quality drop |
| Q2_K | Smallest, most compressed, quality degrades |
Usage
Run with llama.cpp, Ollama, LM Studio, or any GGUF-compatible runtime:
./llama-cli -m NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q4_K_M.gguf -p "Your prompt here"
Notes
- This is a Mixture-of-Experts (A3B) architecture — check RAM/VRAM requirements before choosing a quant.
- For most users, Q4_K_M offers the best balance of speed, size, and output quality.
- Quantized using automated pipeline on Modal with
llama.cpp's conversion and quantization tools.
Credits
- Base model by NVIDIA
- Quantization by NANI-Nithin