Artemis-31B-v1.2 GGUF
Static GGUF quantizations of TheDrummer/Artemis-31B-v1.2.
Artemis-31B-v1.2 is a fine-tuned, unaligned creative writing and roleplay model built on Google's Gemma 4 31B base architecture. It is specifically calibrated for literary depth, psychological nuance, long-horizon scene continuity, and multi-turn roleplay.
Available Files & VRAM Recommendations
| Filename | Size | Recommended VRAM | Description & Use Case |
|---|---|---|---|
Artemis-31B-v1.2-Q3_K_M.gguf |
15.3 GB | 18 – 20 GB | Low-memory footprint; runs comfortably on constrained setups. |
Artemis-31B-v1.2-Q4_K_S.gguf |
17.8 GB | 20 – 22 GB | Lean 4-bit quant; fits tighter single-GPU configurations. |
Artemis-31B-v1.2-Q4_K_M.gguf |
18.7 GB | 24 GB | Community Sweet Spot. Balances quality retention and context headroom. |
Artemis-31B-v1.2-Q5_K_M.gguf |
21.8 GB | 24 – 32 GB | Near-lossless precision; exceptional for descriptive prose and dialogue fidelity. |
Artemis-31B-v1.2-Q6_K.gguf |
25.2 GB | 32 – 40 GB | High-tier quant; indistinguishable from native 16-bit float. |
Artemis-31B-v1.2-Q8_0.gguf |
32.6 GB | 40 – 48 GB | Reference baseline quant with near-zero perplexity loss. |
Prompt Formats
Artemis-31B uses the Gemma 4 Chat Template and natively supports both Thinking (Reasoning Scratchpad) and Direct Storytelling operational modes.
1. Thinking Mode (Recommended for Complex Plots)
Trigger the reasoning planner by embedding <|think|> into the prompt. The model will plan character motivations, spatial distance, and scene tone in a scratchpad before generating the visible narrative.
<start_of_turn>user
<|think|>
Write a scene where two rival mercenaries negotiate a temporary truce in an abandoned chapel during a midnight thunderstorm.<end_of_turn>
<start_of_turn>model
<|channel>thought
The scene demands tension and heavy atmospheric detail. Both characters should remain guarded, tracking the other's weapon hand.
<|channel>call
Rain rattled against the leaded glass like loose teeth...
2. Standard Non-Thinking Mode
For traditional turn-by-turn dialogue without reasoning tokens:
<start_of_turn>user
{{user_prompt}}<end_of_turn>
<start_of_turn>model
{{model_response}}<end_of_turn>
Sampler Guidance (Anti-"Dash Spiral")
Because Artemis possesses a wide, expressive vocabulary distribution, standard greedy samplers can sometimes trigger repetitive loops or excessive em-dashes (—). The community consensus recommends pairing the model with Min-P:
Recommended Daily Driver (SillyTavern / KoboldCpp)
- Temperature:
0.95–1.05 - Min-P:
0.08 - Top-P: Disabled (
1.0) - Repetition Penalty:
1.06–1.08 - Repetition Penalty Range:
2048tokens - Presence / Frequency Penalty:
0.00
Quickstart Guide
Running via llama.cpp
# Serve as a local OpenAI-compatible API endpoint
llama-server \
-hf Abiray/Artemis-31B-v1.2-GGUF:Q4_K_M \
-c 8192 \
-ngl 99 \
--temp 0.95 \
--min-p 0.08 \
--repeat-penalty 1.06 \
--port 8080
Credits & Acknowledgments
- Original Model: TheDrummer for creating and fine-tuning Artemis-31B-v1.2.
- Base Architecture: Google DeepMind's Gemma 4 31B.
- Quantization Engine: llama.cpp by Georgi Gerganov and contributors.