auto-200m-2-int8
This is auto-200m-2 with its weights stored as 8-bit integers. The file is 150 MB instead of 299 MB, and it gets the same benchmark score: 2,890/3,000, with 53 false approvals and 57 false denials. Only 2 of the 3,000 decisions differ from the BF16 model.
auto-200m-2 is a 149.6M-parameter ModernBERT classifier. It reads an AI agent's proposed tool call, the user's request and the agent's history, then answers approve or deny. It takes up to 65,536 tokens of context. This version loads through one small Python file, auto_quant.py, in plain PyTorch, without compiled kernels. It was checked on NVIDIA CUDA, the Apple M4 Max CPU and Apple MPS. For an even smaller file, see auto-200m-2-int4 (77 MB).
Results
These results use the pinned 3,000-item Approve-or-Deny benchmark (revision a38b6259), full input lengths, P(deny) >= 0.5, and the same evaluation code as the base model card. The quantized rows were computed on CUDA in BF16 from the stored integer weights. A false approval is an unsafe call that was approved. A false denial is an authorized call that was denied.
| Model | Weights file | Accuracy | False approvals | False denials | AUROC | 16k–64k tokens | Validation audit | Decisions that differ from BF16 |
|---|---|---|---|---|---|---|---|---|
| auto-200m-2 (BF16) | 299 MB | 96.33% (2890) | 53/1401 | 57/1599 | 0.9937 | 94.14% | 98.34% | — |
| auto-200m-2-int8 | 150 MB | 96.33% (2890) | 53/1401 | 57/1599 | 0.9937 | 94.14% | 98.34% | 2 |
| auto-200m-2-int4 | 77 MB | 96.13% (2884) | 60/1401 | 56/1599 | 0.9933 | 94.98% | 98.15% | 44 |
The 16k–64k column covers 239 benchmark items. The validation audit is 2,595 validation rows that were never used for training or selection. Paired with the BF16 model on the same items, accuracy differs by +0.00 points (95% interval -0.10 to +0.10; 1 items right only for BF16, 1 right only for int8; exact McNemar p = 1.00). On the tool probes it gets 24/24 published skills/MCP/custom-tool probes and 38/40 fresh scope/history/injection probes; the BF16 model gets 24/24 and 38/40. Breakdowns by category, language, difficulty and length are in eval_results.json, and per-item logits are in benchmark_predictions.npz.
Devices
| Device | Compute dtype | Items | Accuracy on those items | Same items, CUDA | Decisions that differ from CUDA | Largest P(deny) difference |
|---|---|---|---|---|---|---|
| NVIDIA RTX PRO 6000 (CUDA) | BF16 | all 3,000 | 96.33% | reference | — | — |
| NVIDIA RTX PRO 6000, public loader (both attention modes) | BF16 | 300 under 2,048 tokens | — | — | 0 | 0.018 |
| Apple M4 Max CPU | FP32 | 2,452 up to 2,048 tokens | 96.82% | 96.86% | 1 | 0.090 |
| Apple M4 Max GPU (MPS) | FP32 | all 3,000 (longest 56,176 tokens) | 96.30% | 96.33% | 1 | 0.090 |
- CUDA row: the benchmark run above.
- Public loader row: 300 benchmark items were scored again with
auto_quant.loadon CUDA, once with memory-linear attention and once with PyTorch SDPA, and compared with the benchmark run. - Mac rows:
auto_quant.loadandscorein FP32 on an M4 Max (macOS). The CPU pass used the 2,452 items up to 2,048 tokens. The MPS pass used all 3,000 items, up to 56,176 tokens.
The P(deny) differences come from BF16 on CUDA versus FP32 on the Mac. AMD ROCm should work, since the loader is plain PyTorch, but it wasn't tested. The details are in eval/mac_verify.json.
Memory and speed on the same Mac, one request at a time. All three use the same memory-linear attention:
| Model | Device | Memory after load | Median, under 1k tokens | 7,449 tokens | 29,233 tokens |
|---|---|---|---|---|---|
| auto-200m-2 (BF16 file, runs in FP32) | CPU | 806 MB | 68 ms | 1.65 s | 15.7 s |
| int8 | CPU | 379 MB | 70 ms | 1.49 s | 13.4 s |
| int4 | CPU | 305 MB | 104 ms | 1.51 s | 12.5 s |
| auto-200m-2 (BF16 file, runs in FP32) | MPS | 570 MB | 24 ms | 0.46 s | 4.4 s |
| int8 | MPS | 143 MB | 30 ms | 0.47 s | 4.1 s |
| int4 | MPS | 74 MB | 38 ms | 0.47 s | 4.3 s |
Each model and device ran in its own process. Latency is the median over 110 short benchmark items. On CPU, memory is how much the process grew during loading. On MPS, it's the tensors held on the GPU. The BF16 checkpoint runs in FP32 on CPU and MPS, which is Transformers' default there.
On the GPU, the int8 and int4 weights take about 4× and 8× less memory than the FP32 copy. They don't make short requests faster, though: each layer turns its integer weights back into floats on every call, and unpacking 4-bit values costs extra, so int4 is the slowest on short inputs. On long inputs, activations dominate both time and memory; at 29,233 tokens the CPU peak was 2.1 / 1.8 / 1.7 GB (BF16 / int8 / int4). See eval/speed_mac.json.
How it was made
- Format. Weights are stored as symmetric int8 in [-127, 127], with one FP16 scale per output channel (per row for the token embedding). The token embeddings, every attention and MLP linear layer, and
head.denseare quantized. The LayerNorm weights and the final 768×2 classifier stay in FP16. Activations keep the device's float type, and each layer rebuilds its own float weights just before its matrix multiply. The details are in quant_config.json and the docstring ofauto_quant.py. - Quantization-aware training (QAT), and why these weights aren't from it. QAT ran on the base model: fake quantization with straight-through rounding, learnable (LSQ) scales, loss
0.1·CE + 0.9·KL(base ‖ quantized)at T = 1, 150,000 training rows and 4 exports. A rule fixed before any benchmark run froze whichever candidate had the lowest loss (NLL) on the 7,824-row validation selection split. Plain post-training quantization (PTQ, MSE-clipped scales) was one of the candidates. PTQ won with an NLL of 0.07195, against 0.07194 for the BF16 model itself. Every QAT export came out slightly higher (0.0730–0.0733). So these weights are the PTQ export. At 8 bits, rounding alone loses almost nothing on this model. All candidates are in eval/qat_candidates.json. - One benchmark look. The frozen export was benchmarked once.
- Base weights. This quantizes the original auto-200m-2 (revision
0bbb929f). In the same job, Auto 3B was distilled into auto-200m-2 as a candidate "iteration 1". It made fewer false approvals but more false denials, so it failed its gate and wasn't published, and both quants start from the original weights.
Usage
import os, sys
from huggingface_hub import hf_hub_download
repo = "ProCreations/auto-200m-2-int8"
sys.path.insert(0, os.path.dirname(hf_hub_download(repo, "auto_quant.py")))
import auto_quant
model, tokenizer = auto_quant.load(repo) # picks CUDA/ROCm, then Apple MPS, then CPU
text = auto_quant.build_input(
user_request="Clean up the build artifacts and reinstall dependencies.",
history=[{"tool": "Bash", "args": "ls", "result": "node_modules dist package.json"}],
call={"tool": "Bash", "args": "rm -rf node_modules dist && npm install"},
)
p_deny = auto_quant.score(model, tokenizer, text)
print("deny" if p_deny >= 0.5 else "approve", round(p_deny, 3))
You need torch, transformers 5.x, huggingface_hub and safetensors; nothing gets compiled. Labels are 0 = approve and 1 = deny. Keep the three section headers exactly as build_input writes them. Arguments are strings; for JSON arguments, compact JSON (json.dumps(args, separators=(",", ":"))) matches the training data.
score() takes one request at a time, up to 65,536 tokens, and its memory grows linearly with length. For padded batches, load with attention="sdpa" and call model(**inputs). AutoModelForSequenceClassification.from_pretrained can't read this format. The Auto runtime 0.2.0 runs the BF16 model and doesn't load these files yet.
Files
auto_quant_int8.safetensors: the quantized weights.quant_config.jsondescribes the format.config.json,tokenizer.json,tokenizer_config.json: from auto-200m-2.auto_quant.py: the loader, about 280 lines of PyTorch.eval_results.jsonandbenchmark_predictions.npz: benchmark, audit and probe results, plus per-item logits.eval/: the PTQ sweep, the QAT candidates and their validation scores, the probe results and the device checks.training/: the QAT and evaluation code. It ran inside a temporary job, so paths refer to that job.
Limitations
This is a classifier, not a policy engine. It approves routine authorized work, and it denies consequential unauthorized actions and actions that follow injected instructions. It can't inspect hidden file contents, resolve opaque executables or know what a URL will do at runtime, so false approvals remain possible. The labels are synthetic, and the benchmark has been reused across Auto releases, so evaluate it on your own traffic before relying on it.