Abiray/Artemis-31B-v1.2-GGUF

🤗 Hugging Face sourceany-to-anyapache-2.031B activated134 GBGGUF✓ 8 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo Abiray/Artemis-31B-v1.2-GGUF ./model-folder
Needs a seeder →

Artemis-31B-v1.2 GGUF

Static GGUF quantizations of TheDrummer/Artemis-31B-v1.2.

Artemis-31B-v1.2 is a fine-tuned, unaligned creative writing and roleplay model built on Google's Gemma 4 31B base architecture. It is specifically calibrated for literary depth, psychological nuance, long-horizon scene continuity, and multi-turn roleplay.


Available Files & VRAM Recommendations

Filename Size Recommended VRAM Description & Use Case
Artemis-31B-v1.2-Q3_K_M.gguf 15.3 GB 18 – 20 GB Low-memory footprint; runs comfortably on constrained setups.
Artemis-31B-v1.2-Q4_K_S.gguf 17.8 GB 20 – 22 GB Lean 4-bit quant; fits tighter single-GPU configurations.
Artemis-31B-v1.2-Q4_K_M.gguf 18.7 GB 24 GB Community Sweet Spot. Balances quality retention and context headroom.
Artemis-31B-v1.2-Q5_K_M.gguf 21.8 GB 24 – 32 GB Near-lossless precision; exceptional for descriptive prose and dialogue fidelity.
Artemis-31B-v1.2-Q6_K.gguf 25.2 GB 32 – 40 GB High-tier quant; indistinguishable from native 16-bit float.
Artemis-31B-v1.2-Q8_0.gguf 32.6 GB 40 – 48 GB Reference baseline quant with near-zero perplexity loss.

Prompt Formats

Artemis-31B uses the Gemma 4 Chat Template and natively supports both Thinking (Reasoning Scratchpad) and Direct Storytelling operational modes.

1. Thinking Mode (Recommended for Complex Plots)

Trigger the reasoning planner by embedding <|think|> into the prompt. The model will plan character motivations, spatial distance, and scene tone in a scratchpad before generating the visible narrative.

<start_of_turn>user
<|think|>
Write a scene where two rival mercenaries negotiate a temporary truce in an abandoned chapel during a midnight thunderstorm.<end_of_turn>
<start_of_turn>model
<|channel>thought
The scene demands tension and heavy atmospheric detail. Both characters should remain guarded, tracking the other's weapon hand.
<|channel>call
Rain rattled against the leaded glass like loose teeth...

2. Standard Non-Thinking Mode

For traditional turn-by-turn dialogue without reasoning tokens:

<start_of_turn>user
{{user_prompt}}<end_of_turn>
<start_of_turn>model
{{model_response}}<end_of_turn>

Sampler Guidance (Anti-"Dash Spiral")

Because Artemis possesses a wide, expressive vocabulary distribution, standard greedy samplers can sometimes trigger repetitive loops or excessive em-dashes (—). The community consensus recommends pairing the model with Min-P:

Recommended Daily Driver (SillyTavern / KoboldCpp)

  • Temperature: 0.95 – 1.05
  • Min-P: 0.08
  • Top-P: Disabled (1.0)
  • Repetition Penalty: 1.06 – 1.08
  • Repetition Penalty Range: 2048 tokens
  • Presence / Frequency Penalty: 0.00

Quickstart Guide

Running via llama.cpp

# Serve as a local OpenAI-compatible API endpoint
llama-server \
  -hf Abiray/Artemis-31B-v1.2-GGUF:Q4_K_M \
  -c 8192 \
  -ngl 99 \
  --temp 0.95 \
  --min-p 0.08 \
  --repeat-penalty 1.06 \
  --port 8080

Credits & Acknowledgments

  • Original Model: TheDrummer for creating and fine-tuning Artemis-31B-v1.2.
  • Base Architecture: Google DeepMind's Gemma 4 31B.
  • Quantization Engine: llama.cpp by Georgi Gerganov and contributors.