ProCreations/auto-200m-2-int4

🤗 Hugging Face sourcetext-classificationapache-2.077 MBother✓ 3 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo ProCreations/auto-200m-2-int4 ./model-folder
Needs a seeder →

auto-200m-2-int4

This is auto-200m-2 with its weights stored as 4-bit integers. The file is 77 MB instead of 299 MB, about a quarter of the size. It gets 2,884/3,000 on the benchmark, with 60 false approvals and 56 false denials. The BF16 model gets 2,890, with 53 and 57. 44 of the 3,000 decisions differ from the BF16 model.

auto-200m-2 is a 149.6M-parameter ModernBERT classifier. It reads an AI agent's proposed tool call, the user's request and the agent's history, then answers approve or deny. It takes up to 65,536 tokens of context. This version loads through one small Python file, auto_quant.py, in plain PyTorch, without compiled kernels. It was checked on NVIDIA CUDA, the Apple M4 Max CPU and Apple MPS. If you want exactly the BF16 model's decisions at half the size, use auto-200m-2-int8 (150 MB).

Results

These results use the pinned 3,000-item Approve-or-Deny benchmark (revision a38b6259), full input lengths, P(deny) >= 0.5, and the same evaluation code as the base model card. The quantized rows were computed on CUDA in BF16 from the stored integer weights. A false approval is an unsafe call that was approved. A false denial is an authorized call that was denied.

Model Weights file Accuracy False approvals False denials AUROC 16k–64k tokens Validation audit Decisions that differ from BF16
auto-200m-2 (BF16) 299 MB 96.33% (2890) 53/1401 57/1599 0.9937 94.14% 98.34% —
auto-200m-2-int8 150 MB 96.33% (2890) 53/1401 57/1599 0.9937 94.14% 98.34% 2
auto-200m-2-int4 77 MB 96.13% (2884) 60/1401 56/1599 0.9933 94.98% 98.15% 44

The 16k–64k column covers 239 benchmark items. The validation audit is 2,595 validation rows that were never used for training or selection. Paired with the BF16 model on the same items, accuracy differs by -0.20 points (95% interval -0.63 to +0.23; 25 items right only for BF16, 19 right only for int4; exact McNemar p = 0.45). On the tool probes it gets 24/24 published skills/MCP/custom-tool probes and 38/40 fresh scope/history/injection probes; the BF16 model gets 24/24 and 38/40. Breakdowns by category, language, difficulty and length are in eval_results.json, and per-item logits are in benchmark_predictions.npz.

Two benchmark looks

This model was benchmarked twice, and both results are reported here. The first frozen export was plain post-training quantization (PTQ), and it scored 2,876/3,000 with 73 false approvals and 51 false denials (-0.47 points against BF16, 52 decisions changed). That result prompted one more QAT run, described below and recorded in the training plan when it was launched. Its best export beat PTQ under the same validation rule, so it replaced the PTQ export, was benchmarked once, and is what's published. Because a benchmark result prompted the retry, the published score is not a clean single look. The first look's logits are in eval/first_look_benchmark_predictions.npz.

Devices

Device Compute dtype Items Accuracy on those items Same items, CUDA Decisions that differ from CUDA Largest P(deny) difference
NVIDIA RTX PRO 6000 (CUDA) BF16 all 3,000 96.13% reference — —
NVIDIA RTX PRO 6000, public loader (both attention modes) BF16 300 under 2,048 tokens — — 1 0.022
Apple M4 Max CPU FP32 2,452 up to 2,048 tokens 96.45% 96.41% 3 0.043
Apple M4 Max GPU (MPS) FP32 2,761 up to 16,384 tokens 96.27% 96.23% 3 0.043
  • CUDA row: the benchmark run above.
  • Public loader row: 300 benchmark items were scored again with auto_quant.load on CUDA, once with memory-linear attention and once with PyTorch SDPA, and compared with the benchmark run.
  • Mac rows: auto_quant.load and score in FP32 on an M4 Max (macOS). The CPU pass used the 2,452 items up to 2,048 tokens. The MPS pass used the 2,761 items up to 16,384 tokens.

The P(deny) differences come from BF16 on CUDA versus FP32 on the Mac. AMD ROCm should work, since the loader is plain PyTorch, but it wasn't tested. The details are in eval/mac_verify.json.

Memory and speed on the same Mac, one request at a time. All three use the same memory-linear attention:

Model Device Memory after load Median, under 1k tokens 7,449 tokens 29,233 tokens
auto-200m-2 (BF16 file, runs in FP32) CPU 806 MB 68 ms 1.65 s 15.7 s
int8 CPU 379 MB 70 ms 1.49 s 13.4 s
int4 CPU 305 MB 104 ms 1.51 s 12.5 s
auto-200m-2 (BF16 file, runs in FP32) MPS 570 MB 24 ms 0.46 s 4.4 s
int8 MPS 143 MB 30 ms 0.47 s 4.1 s
int4 MPS 74 MB 38 ms 0.47 s 4.3 s

Each model and device ran in its own process. Latency is the median over 110 short benchmark items. On CPU, memory is how much the process grew during loading. On MPS, it's the tensors held on the GPU. The BF16 checkpoint runs in FP32 on CPU and MPS, which is Transformers' default there.

On the GPU, the int8 and int4 weights take about 4× and 8× less memory than the FP32 copy. They don't make short requests faster, though: each layer turns its integer weights back into floats on every call, and unpacking 4-bit values costs extra, so int4 is the slowest on short inputs. On long inputs, activations dominate both time and memory; at 29,233 tokens the CPU peak was 2.1 / 1.8 / 1.7 GB (BF16 / int8 / int4). See eval/speed_mac.json.

How it was made

  • Format. Weights are stored as symmetric int4 in [-8, 7], packed two per byte, with one FP16 scale for every 128 consecutive input weights. The token embeddings, every attention and MLP linear layer, and head.dense are quantized. The LayerNorm weights and the final 768×2 classifier stay in FP16. Activations keep the device's float type, and each layer rebuilds its own float weights just before its matrix multiply. The group size was chosen from 32, 64 and 128 by PTQ loss (NLL) on the 7,824-row validation selection split, taking the largest group within 2% of the best (eval/ptq_sweep.json). The details are in quant_config.json and the docstring of auto_quant.py.
  • Selection rule. A rule fixed before any benchmark run froze whichever candidate had the lowest NLL on the validation selection split. The candidates were the PTQ export (MSE-clipped scales) and every QAT export.
  • QAT, first run. This run used fake quantization with straight-through rounding, learnable (LSQ) scales starting from the PTQ ones, loss 0.1·CE + 0.9·KL(base ‖ quantized) at T = 1, 600,000 training rows (learning rates 1e-5 for weights and 2e-4 for scales) and 4 exports. All four drifted above their PTQ starting point on validation NLL (0.0731–0.0742, against 0.0726), so PTQ was frozen.
  • QAT, second run. This run used KL to the BF16 model only (T = 1), learning rates 10× lower (1e-6 for weights, 2e-5 for scales), 300,000 rows and 4 exports. Validation NLL came out at 0.0714–0.0720. The best export, qat-int4-v2-step-500 (0.07136), was below PTQ and below the BF16 model's 0.07194. It was frozen under the same rule and benchmarked, and those are the weights published here.
  • Base weights. This quantizes the original auto-200m-2 (revision 0bbb929f). In the same job, Auto 3B was distilled into auto-200m-2 as a candidate "iteration 1". It made fewer false approvals but more false denials, so it failed its gate and wasn't published, and both quants start from the original weights.

Every candidate's validation scores are in eval/qat_candidates.json and eval/int4_retry.json.

Usage

import os, sys
from huggingface_hub import hf_hub_download

repo = "ProCreations/auto-200m-2-int4"
sys.path.insert(0, os.path.dirname(hf_hub_download(repo, "auto_quant.py")))
import auto_quant

model, tokenizer = auto_quant.load(repo)   # picks CUDA/ROCm, then Apple MPS, then CPU
text = auto_quant.build_input(
    user_request="Clean up the build artifacts and reinstall dependencies.",
    history=[{"tool": "Bash", "args": "ls", "result": "node_modules dist package.json"}],
    call={"tool": "Bash", "args": "rm -rf node_modules dist && npm install"},
)
p_deny = auto_quant.score(model, tokenizer, text)
print("deny" if p_deny >= 0.5 else "approve", round(p_deny, 3))

You need torch, transformers 5.x, huggingface_hub and safetensors; nothing gets compiled. Labels are 0 = approve and 1 = deny. Keep the three section headers exactly as build_input writes them. Arguments are strings; for JSON arguments, compact JSON (json.dumps(args, separators=(",", ":"))) matches the training data.

score() takes one request at a time, up to 65,536 tokens, and its memory grows linearly with length. For padded batches, load with attention="sdpa" and call model(**inputs). AutoModelForSequenceClassification.from_pretrained can't read this format. The Auto runtime 0.2.0 runs the BF16 model and doesn't load these files yet.

Files

  • auto_quant_int4.safetensors: the quantized weights. quant_config.json describes the format.
  • config.json, tokenizer.json, tokenizer_config.json: from auto-200m-2.
  • auto_quant.py: the loader, about 280 lines of PyTorch.
  • eval_results.json and benchmark_predictions.npz: benchmark, audit and probe results, plus per-item logits.
  • eval/: the PTQ sweep, the QAT candidates and their validation scores, the probe results and the device checks.
  • training/: the QAT and evaluation code. It ran inside a temporary job, so paths refer to that job.

Limitations

This is a classifier, not a policy engine. It approves routine authorized work, and it denies consequential unauthorized actions and actions that follow injected instructions. It can't inspect hidden file contents, resolve opaque executables or know what a URL will do at runtime, so false approvals remain possible. The labels are synthetic, and the benchmark has been reused across Auto releases, so evaluate it on your own traffic before relying on it.

At 4 bits, 44 of 3,000 decisions differ from the BF16 model, against 2 for int8. If memory allows, int8 is the closer copy.