mys/plaidq-0.7b-16step-GGUF

🤗 Hugging Face sourceapache-2.0700M activated2.9 GBGGUF✓ 3 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo mys/plaidq-0.7b-16step-GGUF ./model-folder
Needs a seeder →

PlaidQ 0.7B 16-step GGUF (ggmlc)

Continuous latent diffusion code tab-completion model compiled from fredzzp/plaidq-0.7b-16step (Qwen3-0.6B trunk, codebook dim 16, 16-step DDIM).

These files are not llama.cpp / llama-cli GGUFs. They are produced by ggmlc, a neural network compiler that lowers PyTorch / JAX / Flax / Keras models to high-performance GGML execution. Loading them in llama.cpp will fail.

Get running

  1. Download a GGUF from this repo (see table below).
  2. Download the pre-built tab_completion binary from ggmlc GitHub Releases (or all releases).
  3. CLI / daemon details — see the example README: examples/tab_completion (raw README).

The host-side DDIM sampler, FIM canvas, and codebook live in that standalone C++ binary — not in llama.cpp.

Files

File Quant Size (approx) Notes
plaidq_0.7b_16step_f16.gguf F16 ~1.4 GB Reference / highest quality
plaidq_0.7b_16step_q8_0.gguf Q8_0 ~740 MB Good quality, smaller VRAM
plaidq_0.7b_16step_ud_q4_k_m.gguf UD_Q4_K_M ~600 MB Unsloth-dynamic Q4_K_M; 1D norms stay F32

Each GGUF embeds the Qwen3 BPE tokenizer, the embedding_matrix codebook, and learned noise-schedule knots (plaidq.schedule_t / plaidq.schedule_g).

huggingface-cli download mys/plaidq-0.7b-16step-GGUF plaidq_0.7b_16step_ud_q4_k_m.gguf --local-dir .

Quick start

Put the release tab_completion binary on your PATH (or run it by path), then:

# 16-step completion on CUDA
.\tab_completion.exe complete `
  plaidq_0.7b_16step_ud_q4_k_m.gguf `
  --prefix "def add(a, b):`n    " `
  --steps 16 --score-temp 0.5 --max-tokens 32 --device cuda
./tab_completion complete \
  plaidq_0.7b_16step_ud_q4_k_m.gguf \
  --prefix $'def add(a, b):\n    ' \
  --steps 16 --score-temp 0.5 --max-tokens 32 --device cuda

Example smoke output (UD_Q4_K_M, RTX 4050 Laptop):

 return a + b

# Test the function
print(add(5, 7))  # Output: 12

Notes

  • Canvas length is fixed at 256 tokens (CUDA-graph friendly).
  • Use --steps 16 with --score-temp 0.5 for the distilled 16-step checkpoint.
  • The runtime forces FP32 cuBLAS accumulate on CUDA (GGML_CUDA_CUBLAS_COMPUTE_TYPE=f32) so the vocab head stays finite on Ada GPUs.

License

Apache 2.0, same as fredzzp/plaidq-0.7b-16step. Compiler: ggmlc (MIT).