distil-labs/Distil-PII-SmolLM2-135M-Instruct

🤗 Hugging Face sourcetext-generationapache-2.0135M params269 MBsafetensors✓ 1 checksumupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo distil-labs/Distil-PII-SmolLM2-135M-Instruct ./model-folder
Needs a seeder →

Distil-PII-SmolLM2-135M-Instruct

A small language model (SLM) fine-tuned by Distil Labs for policy-aware PII redaction that outputs a single JSON object with redacted_text and entities. Optimized to run locally with strong accuracy and strict schema adherence.

Model Details

  • Developed by: Distil Labs GmbH
  • License: Apache 2
  • Finetuned from: HuggingFaceTB/SmolLM2-135M-Instruct

Intended Use & Limitations

  • Use cases: Redacting support chats, logs, tickets, transcripts—removing identity while preserving ops signals (IDs last-4, order numbers, etc.).
  • Out of scope: Legal or compliance advice; languages beyond English (generalization not guaranteed); domain-specific IDs unseen in training.

Input & Output

Input: A plain-text prompt with task instruction + context. Output (JSON only):

{
  "redacted_text": "Text with in-place tokens",
  "entities": [
    {"value": "<original>", "replacement_token": "[TOKEN]", "reason": "<why>"}
  ]
}

Tokens: [PERSON] [EMAIL] [PHONE] [ADDRESS] [SSN] [ID] [UUID] [CARD_LAST4:####] [IBAN_LAST4:####] [GENDER] [AGE] [RACE] [MARITAL_STATUS]

Training

Instruction-tuned on a compact policy spec + ~20 curated examples emphasizing exact JSON schema, minimal in-place edits, and entity correctness.

Evaluation

Judged by a frontier LLM using a deterministic rubric: JSON-only, schema validity, redacted_text exact match, and set-equality of (value, replacement_token) pairs (reason/order ignored). Score: 0.25 +/- 0.05.

How to Use

Details of deployment can be found in https://docs.distillabs.ai/how-to/model-deployment

Risks & Mitigations

  • False negatives/positives: May miss novel formats or over-redact generic terms. Mitigate via guardrails + post-validation.
  • Policy drift: Keep task preamble fixed; monitor with unit tests.

Model Sources