Abiray/Artemis-31B-v1.2-GGUF

🤗 Hugging Face 来源any-to-anyapache-2.0激活 31B134 GBGGUF✓ 8 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo Abiray/Artemis-31B-v1.2-GGUF ./model-folder
需要做种者 →

Artemis-31B-v1.2 GGUF

Static GGUF quantizations of TheDrummer/Artemis-31B-v1.2.

Artemis-31B-v1.2 is a fine-tuned, unaligned creative writing and roleplay model built on Google's Gemma 4 31B base architecture. It is specifically calibrated for literary depth, psychological nuance, long-horizon scene continuity, and multi-turn roleplay.


Available Files & VRAM Recommendations

Filename Size Recommended VRAM Description & Use Case
Artemis-31B-v1.2-Q3_K_M.gguf 15.3 GB 18 – 20 GB Low-memory footprint; runs comfortably on constrained setups.
Artemis-31B-v1.2-Q4_K_S.gguf 17.8 GB 20 – 22 GB Lean 4-bit quant; fits tighter single-GPU configurations.
Artemis-31B-v1.2-Q4_K_M.gguf 18.7 GB 24 GB Community Sweet Spot. Balances quality retention and context headroom.
Artemis-31B-v1.2-Q5_K_M.gguf 21.8 GB 24 – 32 GB Near-lossless precision; exceptional for descriptive prose and dialogue fidelity.
Artemis-31B-v1.2-Q6_K.gguf 25.2 GB 32 – 40 GB High-tier quant; indistinguishable from native 16-bit float.
Artemis-31B-v1.2-Q8_0.gguf 32.6 GB 40 – 48 GB Reference baseline quant with near-zero perplexity loss.

Prompt Formats

Artemis-31B uses the Gemma 4 Chat Template and natively supports both Thinking (Reasoning Scratchpad) and Direct Storytelling operational modes.

1. Thinking Mode (Recommended for Complex Plots)

Trigger the reasoning planner by embedding <|think|> into the prompt. The model will plan character motivations, spatial distance, and scene tone in a scratchpad before generating the visible narrative.

<start_of_turn>user
<|think|>
Write a scene where two rival mercenaries negotiate a temporary truce in an abandoned chapel during a midnight thunderstorm.<end_of_turn>
<start_of_turn>model
<|channel>thought
The scene demands tension and heavy atmospheric detail. Both characters should remain guarded, tracking the other's weapon hand.
<|channel>call
Rain rattled against the leaded glass like loose teeth...

2. Standard Non-Thinking Mode

For traditional turn-by-turn dialogue without reasoning tokens:

<start_of_turn>user
{{user_prompt}}<end_of_turn>
<start_of_turn>model
{{model_response}}<end_of_turn>

Sampler Guidance (Anti-"Dash Spiral")

Because Artemis possesses a wide, expressive vocabulary distribution, standard greedy samplers can sometimes trigger repetitive loops or excessive em-dashes (—). The community consensus recommends pairing the model with Min-P:

Recommended Daily Driver (SillyTavern / KoboldCpp)

  • Temperature: 0.95 – 1.05
  • Min-P: 0.08
  • Top-P: Disabled (1.0)
  • Repetition Penalty: 1.06 – 1.08
  • Repetition Penalty Range: 2048 tokens
  • Presence / Frequency Penalty: 0.00

Quickstart Guide

Running via llama.cpp

# Serve as a local OpenAI-compatible API endpoint
llama-server \
  -hf Abiray/Artemis-31B-v1.2-GGUF:Q4_K_M \
  -c 8192 \
  -ngl 99 \
  --temp 0.95 \
  --min-p 0.08 \
  --repeat-penalty 1.06 \
  --port 8080

Credits & Acknowledgments

  • Original Model: TheDrummer for creating and fine-tuning Artemis-31B-v1.2.
  • Base Architecture: Google DeepMind's Gemma 4 31B.
  • Quantization Engine: llama.cpp by Georgi Gerganov and contributors.